Abstract
Video data analysis (VDA) represents an important methodological framework for contemporary research approaches to the myriad of footage available from cameras, devices, and phones. Footage from police body-worn cameras (BWCs) is anticipated to be a widely available platform for social science researchers to scrutinize the interactions between police and citizens. We examine issues of validity and reliability as related to BWCs in the context of VDA, based on an assessment of the quality of audio and video obtained from that platform. Second, we compare the coding of BWC footage obtained from a sample of police-citizen encounters to coding of the same events by on-scene coders using an instrument adapted from in-person systematic social observations (SSOs). Findings show that there are substantial and systematic audio and video gaps present in BWC footage as a source of data for social science investigation that likely impact the reliability of measures. Despite these problems, BWC data have substantial capacity for judging sequential developments, causal ordering, and the duration of events. Thus, the technology should open theoretical frames that are too cumbersome for in-person observation. Theoretical development with VDA in mind is suggested as an important pathway for future researchers in terms of framing data collection from BWCs and also suggesting areas where triangulation is essential.
Keywords
Introduction
Body-worn camera (BWC) technology has emerged as a vital tool for transparency, accountability, and evidentiary concerns in American and international policing agencies. BWC footage also represents a data source ripe for video data analysis (VDA) applications that comport with systematic social observation (SSO) (Sytsma, Chillar, and Piza 2021). Reiss (1971) pioneered in-person SSO (IPSSO) of police in his study of Boston, Washington D.C., and Chicago departments. This offered insights into the interactions between police and citizens while yielding novel accounts of disrespect and coercion by social control agents. Importantly, direct observations of police activity through this method enabled researchers to chronicle actual and real evidence of police behavior that went beyond administrative data (i.e., incident reports, calls for service, and arrests) and outcome-related data (i.e., uses of force and civilian complaints). Because of the labor-intensive nature of observation and logistical difficulties of getting access to police agencies and then riding along with police officers, systematic observation of police has been a relatively rare occurrence in police research. However, it has served as an invaluable tool for developing a sense, especially in the United States, about how police deal with citizens. With the proliferation of BWC footage among police departments, however, these data are now poised to offer new insights into police behavior (Worden, McLean, and Bonner 2015).
Despite the large number of BWC research studies and evaluations that have occurred (Lum et al. 2015 and White and Malm 2020), few have examined the challenges, strengths, and the limitations of BWC footage as a source of data and how it can be used to advance inquiry in sociology and psychology frameworks (c.f. Willits and Makin 2018 and Friis et al, 2020). The current study fills a portion of that gap by examining BWC footage from the Los Angeles Police Department (LAPD). We seek to raise practical issues and address a series of research questions that will assist researchers as they contemplate the use of BWC footage in VDA. In particular, we examine how observations from BWC footage compare to existing measures derived from IPSSO. To do so, we make use of observations of LAPD officers in the field and offer a second type referred to as “video-based systematic social observations” or VBSSO from their BWCs. While both of these methods have merit, we need to know more about the reliability and validity of BWC footage in the VDA context. By doing so, we hope to assist sociologists, criminologists, and others as they begin to access and analyze these data.
This article is divided into four major areas. First, we introduce the VDA framework as proposed and described by Nassauer and Legewie (2019; 2021) and then provide background about BWCs in terms of their adoption and their use as data. We briefly discuss the issues for collecting, storing, accessing, and retaining the data by law enforcement agencies. In VDA, BWC Footage, and Technology section, we provide an overview of the literature on BWCs and IPSSOs and VBSSOs and describe how VDA interrelates with the reliability of SSOs. In Literature Review: Using BWC Footage and SSO Research section, we turn our attention to the methods, analysis, and findings of our work with BWC footage and SSOs in Los Angeles. In Discussion and Research Implications section, we discuss the implications of our findings for future research.
VDA, BWC Footage, and Technology
In two articles and a textbook, Nassauer and Legewie (2019; 2021; 2022) propose three important dimensions for assessing validity in VDA, including (1) neutral and balanced data; (2) optimal capture; and (3) natural behavior. The definition of validity they adopt, and to which we refer, is a general one related to these criteria, and whether “concepts and indicators capture what they are intended to capture in the data, and whether the data can sustain the conclusions presented by the researcher” (Nassauer and Legewie 2022:116). For example, neutral and balanced data refer to sources of information that are free from bias. Nassauer and Legewie suggest that researchers consider what types of bias may exist in the data, how bias may affect the analysis, and what can be done to mitigate the bias. Optimal capture means that footage should cover the length of the event, all of the people involved, and the space of the event. Optimal capture asks, how complete or incomplete is the video footage? Natural behavior refers to the degree in which people behave when the camera is on. Do they act differently when being filmed than when the camera is off?
To improve the validity of VDA, Nassauer and Legewie suggest data triangulation. That is, researchers should gather diverse data to overcome the imbalance. Second, triangulating individual data pieces from different sources can address incompleteness. For example, the same scene from different cameras (CCTV or BWC) would assist in this type of event. Lastly, obtaining information from different source materials (e.g., documents, field notes, and participant observation) assists in addressing bias and missing information, and establishes context to inform VDA (Nassauer and Legewie 2021).
How do we apply these concepts to BWC footage created by law enforcement? How and why do police use BWC footage? What is the adequacy of BWC footage from the research and methodological viewpoints of reliability and validity? How can BWC data supplement or enhance other data to measure police behavior and performance?
BWC Technology and Footage
Law enforcement agencies in the United States have adopted BWCs rapidly over the last 7 years (Suss et al. 2018). Following the deaths of Michael Brown in Ferguson, Missouri and Laquan McDonald in Chicago, Illinois, the White House released its report of the President's Task Force on 21st Century Policing (2015), recommending the widespread adoption of police BWCs in the United States (cited in Crow et al. 2017). Shortly thereafter, the United States Department of Justice began to provide funding for law enforcement agencies to purchase BWCs. In 2020, as a result of the George Floyd murder by police officers in Minneapolis, Minnesota, state legislatures began to mandate the use of BWCs by police (e.g., New Jersey, Ohio, and Illinois) as well as the release of video footage to the public (e.g., California). These measures by Federal, state, and local legislatures were part of the drive to make police more accountable and transparent regarding their actions. In addition, BWCs are suggested to generate a civilizing effect on police when interacting with citizens. That is, their self-awareness is increased because they are on video and so they behave in a more polite manner (Boivin et al. 2017).
While the primary intent of BWCs is to improve police behavior, their existence as a data collection tool cannot be ignored. As with other types of video (closed circuit, phone-based, and private surveillance footage), researchers must acknowledge and confront significant limitations that accompany these data. Furthermore, as a data collection tool, BWCs should be viewed in the same light as other police-related datasets like calls for service, crime incidents, arrests, traffic stops, and field interviews.
Limitations of BWC Footage
Before commencing VDA applications, researchers should recognize the limitations of video at the outset. In accessing the video, researchers are likely to most easily obtain data from police department YouTube channels and perhaps from existing departmental BWC archives or databases (Uchida et al. 2022). The limits of YouTube releases include the unknown editing and redaction that is likely to take place. Rarely are raw, unedited videos released to the public, thus researchers must attend to blurring of faces, cuts, and omitted audio as possibilities in that avenue of access. Those sampling and using data from such publicly available platforms must recognize the need to address issues of neutrality and balance. Specifically, agencies uploading data may not be representative of police in general, and the uploaded elements may, similarly, have bias in curation in terms of presentation and included footage. Critical incidents posted on YouTube by LAPD in 2019, for example, represent 36 extreme incidents selected from among the 4 million videos the agency stored in 2019 and which were not shared in whole or in part with the public (Uchida and Anderson 2021). Additionally, the more general limitations outlined below regarding BWC footage would apply to such uploaded video.
Police departments that are willing to share video with researchers may be inclined to pick cases that are believed to be interesting for the question at hand, but this presents problems of generalizability depending on the research question posed and might approximate a convenience sample or sampling on the dependent variable. Even if researchers can access police data directly and randomly sample videos based on traffic stops or other encounters enumerated in police record management systems (RMS) there are still at least four major problems. First, storage space is an issue requiring some departments to routinely delete video to maintain storage and reduce costs (Solomon, Connor and Schmitz 2021). Police departments have different retention policies for video footage, and in some cases, the footage is automatically purged from the system. This may make some sampled cases unavailable, especially those which do not include an arrest, crime, or use of force, which will trigger video retention. Second, video files may be excluded from view because of ongoing investigations. Sytsma, Chillar and Piza (2021) could not access video for specific cases because internal policies prohibited their viewing of “live” and active investigations. Third, video files that are associated with a sampled event may be missing or “untagged” and hard or impossible to locate. This could be due to officers’ not categorizing their footage correctly, a mismatch within the department's data system between elements and identifiers associated with the video, and other factors (Uchida and Anderson 2021). The fourth problem involves the failure to record a video due to non-activation of a sampled event. In some cases, officers do not turn on their cameras (inadvertently or otherwise), though they may write up an incident or arrest report. Recent research by Katz and Huff (2022), using a sample of Phoenix Police Department data, suggests that non-activation of BWCs is substantial and systematically associated with departmental policies as well as a variety of individual, neighborhood, and situational factors. These gaps speak to important validity issues and suggest that the mundane and currently investigated cases may be systematically missing from sample frames.
Once data are obtained, video and audio quality are often cited as barriers to analysis. In addition, researchers have documented the placement of the BWC as problematic when viewing the footage because they are typically mounted on an officer's chest area and clothing and can be dislodged during scuffles or become obstructed by an officer's arms in some tactical positions (Suss et al. 2018; Terrill and Zimmerman 2022). Moreover, the angle of the body camera cannot capture everything that is occurring in front of an officer and does not really capture video of the officer at all. Mateescu, Rosenblat, and Boyd (2016) concluded that even slight movements by the officer can impact what the BWC captures, and the fisheye panoramic view is not reflective of what the wearer is seeing and thus assessing whether stimuli provoke certain responses could be problematic (Boivin et al. 2017). Since the BWC, if activated, “goes where the action is” it represents a set of tradeoffs between perspective and accessibility (Grimshaw 1982).
The limitations and cautions noted above are recognized, to a varying degree, across the growing body of research using directly accessed BWC footage (see Appendix 1). The initial VBSSO-based research includes Mell's (2016) unpublished dissertation coding procedural justice measures from BWC and Voigt et al.'s (2017) research analyzing police-citizen courtesy using the audio track of BWC during traffic stops. Since those two projects at least seven more have drawn on agency BWC footage, coded it, and published peer-reviewed research, with three projects appearing in 2021. Given the likely growth in use of BWC going forward, a systematic comparison of data derived from this source and IPSSO fills an important gap in the methodological literature.
Finally, the ethical framework from which BWC footage is accessed must be carefully weighed by researchers. The principles of beneficence, respect for persons, justice, and respect for law and public interest apply to the access and use of footage for research purposes. For example, to answer some research questions, the use of footage within a private residence may be prohibited whereas review of BWC from public encounters may be available, due to the concern with privacy and lack of consent of citizens balanced against the benefit of answering the particular research question. Further, interstate and international variations in the protection of individual privacy may supersede research interests. Nassauer and Legewie (2022, Chapter 4) offer a general outline of the balance that should be considered in terms of the benefits of the video usage and the uniqueness of the data collected, analysis performed, and the cost-effectiveness of the research overall against the disadvantages and risks to officers and citizens who are involved.
Literature Review: Using BWC Footage and SSO Research
Researchers who have obtained footage have used “video-based systematic social observation” (VBSSO) to explore different ways to measure force and coercion, citizen behaviors, and procedural justice. This section examines the research studies that use VBSSO, their acknowledgment of the problems with the data, and how they fit within the VDA context.
Willits and Makin (2018) analyzed use of force incidents by examining and coding BWC footage from agencies within the state of Washington. They cited several barriers when coding, including the ability to identify or classify the force used by police, the time at which force was used, the angle of the camera during citizen contact, and poor audio quality. Obstacles were apparent when transcribing and coding audio in heavily crowded places due to noise constraints. To overcome these challenges, the researchers examined multiple videos from different officers’ body cameras but still had to reduce their sample size from 174 to 95 videos to include only those meeting their analytical criteria. Sytsma, Chillar and Piza (2021) examined BWC in 122 use-of-force events in Newark, New Jersey and after excluding cases with no camera footage, data unavailable because of an ongoing investigation, incorrect tagging, and other issues, a total of 91 cases were available for analysis. Thus, even in the video which has long retention periods (use of force in these two research projects), the video loss resulted in 25 percent to 45 percent of cases being excluded, reflecting the VDA concern for balanced and neutral data. Specifically, some videos were missing due to non-activation or administrative restriction, others were discarded due to a lack of optimal capture (i.e., the video was not able to speak to the research question at hand). Friis et al. (2020) drew on 123 clips of ticket inspector BWC footage and examined how the micro-interaction of passenger and inspector escalated into aggressive outcomes. The authors report having 374 recordings at the outset but excluded many due to limited capture, poor technical quality of video or audio, or limited duration. In contrast, Terrill (2003) researched patterns of police use of force and citizen resistance using IPSSO among 3,544 suspects sequencing encounters into smaller units and found 83 cases of officers using pain compliance, pushing or throwing citizens, or striking suspects. This suggests that BWC has a substantial capacity for illuminating rare events in new and cost-effective ways, especially with regard to specific sequencing of actions and reactions (Friis et al. 2020; Schafer, Hibdon, and Kyle 2022).
Research using VBSSO has explored the emotional states of officers and suspects from data in 287 police-suspect encounters coded from BWC (Makin et al. 2019). That research reports no data regarding explicit audio or video problems in the sample videos. The authors do, however, recognize that the body cues indicative of negative emotional states of officers may have more limited capture, due to BWC mounting, compared to citizen cues within the video. This implies that the BWC may have differential reliability when measuring similar officer and citizen behaviors or actions. A robust tiered approach to ascertaining inter-rater reliability for that VBSSO, with the aforementioned concerns acknowledged, and a call for further investigation of the reliability concern is suggested (Makin et al. 2019). Holladay and Makin (2021) have recently explored civility of police and citizens in one-on-one encounters, measuring slurs, profanity, deception, and interruptions in 152 cases selected purposefully for having only one officer and one citizen present. One of the drawbacks of this study is that information on the universe of encounters from which the sample is drawn and missing data is not reported. Examples of IPSSO that parallel these questions include Reisig et al. (2004) who used a sample of 3,128 encounters in St. Petersburg, Florida, and Indianapolis, Indiana to explore determinants of citizen disrespect toward police.
Dai (2021) launched a research project in Norfolk, Virginia focused on procedural justice and generated some results from a large sample using BWC focused on 659 randomly selected shifts and 2,861 encounters. No data were reported regarding the quality or missingness of audio or video analyzed for purposes of coding. In parallel, a recent IPSSO examined predictors of procedural justice in a sample of LAPD encounters using 555 citizens to examine the determinants of that dimension of police service (McCluskey et al. 2019).
The VDA Context and SSOs
The comparison of VBSSO to IPSSO, and threats to reliability and validity, can be considered along the three dimensions that Nassauer and Legewie (2019 and 2021) have outlined. For example, the criterion of natural behavior is parallel to that of reactivity in IPSSO. Spano's (2003, 2007) research on IPSSO methodology has, for example, surfaced significant variation in what is observed based on observer gender and time in the field. The impacts he assessed were small but statistically dependable. In comparison, the VDA of BWC lacks a reactivity component that is performative or muted (by the officer impressing an in-person observer or avoiding certain activities) and it is not limited to the observer's single interpretive reference point. That is, multiple persons can view and code video and researchers can determine if codes are reliably applied (Makin, Willits and Brooks 2020) or need to be discarded (Schafer, Hibdon and Kyle 2022). Although the BWC lacks the intrusiveness of in-person observation and has the strength of multiple reviews at lower cost, it does not eliminate or solve the problems that have been raised involving IPSSO research (Spano 2007). For example, mounting of the camera on a participant in the police-citizen encounter might be ascribed as a systematic error or threat to validity of measuring aspects of the police-citizen encounter. This could be akin to that of “over identification” with a research subject's perspective which can be present to a degree within the individual observer. Grimshaw (1982) argues that, for the purposes of addressing some research questions, the type of mount described here can yield important data as it will track officer movements and perhaps be a proxy for attention and focus. However, since mounting is built in to BWC data collection, it is unseen by the video observer, instead perhaps being accepted as “the way things are.” In the context of VDA of BWC, we would frame the concerns regarding data validity and reliability by exploring the extent to which they are not neutral and have suboptimal capture (Nassauer and Legewie 2021).
To explicitly assess the validity and reliability of BWC, we need to consider separate audio and visual channels from which data can be coded. For example, if we are coding physical actions or emotional responses between police and citizen, the process is limited by camera movement and obstruction (Makin et al. 2019). Conversely, the data available to assess officer behavior is largely restricted to verbalization of commands, tone, word choices, and the like, since visual data cannot be ascertained apart from hand movement, gun display, and perhaps the rapidity of body movements. The second officer on the scene, however, would afford opportunities to capture some lead officer behavior and bridge this concern with validity of BWC data. Suspect, victim, or witness behavior may have higher validity in the visual channel since the camera is focused on them, although imperfectly. A complication for VBSSO is establishing a consistent timeline for multiple officer videos, however, timestamps and synchronization may afford reviewers opportunities, with much time cost, for developing accurate timelines and establishing temporal order. Chest-mounted microphones, conversely, likely offer a better record of officer compared to citizen utterances. In sum, we surmise that we would observe differential reliability and validity across these two channels of data that can be drawn from BWC, which may impact the reliability of some common measures of behaviors measured in police-citizen encounters.
Methods, Analysis, and Findings
This section describes the research methods for conducting SSO within the LAPD, how we compared in-person SSOs to video-based SSOs, how we analyzed the data, and the results of the analyses.
The Site
The LAPD serves the second largest city in the United States, Los Angeles, California. With 9,500 sworn personnel and over 2,800 civilian employees, the Department covers the city's 472 square miles. In 2015 the LAPD began its initial deployment of BWCs. Two divisions, Mission and Newton began implementation in September 2015 and are the focus of this study as they were the first two patrol divisions in the Department to receive BWCs. They were specifically selected based on their locations and the differences in their call types and activities. Details regarding both divisions and the overall methodology and the purpose of the larger SSO study are available elsewhere (McCluskey et al. 2019).
Methods
In June 2015, eight observers were trained in a classroom setting on the IPSSO instruments, including discussions of coding protocols, group viewing of vignettes, and a series of field training rides consistent with protocols used in prior studies (Mastrofski et al. 1998). After finalizing all procedures and instruments, observers conducted the initial SSOs in August (Mission Division) and September (Newton Division) of 2015 prior to BWC deployment. Officers were randomly selected for observation in both divisions. The same process was followed in June 2016 for the second wave of observations, after BWC had been deployed for approximately 9 months.
To select the random sample of officers (and up to five alternates) for participation in the SSOs, staff secured a list of all eligible patrol officers with the cooperation of Captains at each Division. The randomly selected officer was identified prior to roll call by the research team (and without input from police personnel to avoid bias). Observers attended division roll calls for all observation shifts and notified the selected officers about the ride-along. For each 6-hour observation period, staff observed the interactions between the assigned officer (O1) and the citizens during each encounter. As LAPD employs two-person patrol cars, each SSO included the randomly selected primary officer (O1) and his/her partner (O2) for that shift. The primary officer observed in the encounter refers to the officer who took the lead in the decision making and had the most interaction with citizens. LAPD tactics, such as assigned roles in “contact and cover” allowed for observers to accurately determine which officer would take the decision making lead prior to the commencement of citizen contact.
Throughout both waves of the SSOs, observers spent 725 h riding with and collecting observational data on the encounters between officers and citizens. A total of 124 ride-alongs (71 from Wave I and 53 from Wave II) were completed for both Newton and Mission Divisions. Completed Wave II rides were then randomly sampled to obtain BWC video from the observed officers to assess and compare data from 102 videos involving 53 events at which IPSSO personnel also collected data. Seven non-encounter events with a total of 13 videos recorded by Officer 1 (O1) and Officer 2 (O2), and 46 citizen encounter events comprising 89 videos were sampled. Duration and distribution of videos across events and encounters are presented in Table 1. VBSSO involved the review of unenhanced video and assessment of the audio and video clarity for each of the 102 videos and coded responses using parallel measures from IPSSO as well as unique questions describing the quality of audio and video rendered by BWC.
Duration and Content of Body-Worn Camera (BWC) Videos Reviewed.
For video reviews, all were conducted on LAPD workstations using a streaming feed to standard desktop monitors. Headphones were used to evaluate the audio stream to isolate reviewers from office background noise. Time spent answering questions on the form ranged from under five minutes (n = 14 videos) to more than an hour (n = 5 videos), with the average time spent being almost 20 min (x̄ = 19.7; SD = 24.7) across the 102 videos (examples of video quality questions assessed by reviewers can be found in Appendix 2).
Analysis and Findings
The series of analyses that follows first examines the BWC audio channel and then assesses the video stream separately to appraise the quality of each and includes the VBSSO assessment of the causes of degradation, noise, or interference in each of the signal channels.
Reliability of the BWC Record
As noted, the reliability of audio and video is unlikely to be equal across officer and non-officer participants (suspects, victims, other first responders, and witnesses) present due to camera and microphone placement. We anticipate, from the extant literature noted by BWC users and coders, that citizen audio will be less reliable when compared to officer audio.
Overall Audio Clarity
Observers conducting video SSO assessed the number of participants’ speech that was audible in the unenhanced BWC reviewed. Participants are defined as officers other than the recording officer, citizens, and other service providers (e.g., emergency medical personnel and other first responders) who are captured by the BWC. As shown in Table 2, the full content of participants’ speech was audible in 19 videos, and most of participants’ content was audible in 45 videos, indicating 63 percent of videos featured full or most of participants’ speech. A separate evaluation of the audio quality of the recording officer, upon which the microphone/camera was mounted was made. Using the same categorical assessment (also in Table 2, right-hand panel) observers coded the BWC footage for the audio clarity of the recording officer's speech in the 98 videos where the recording officer spoke. In 45 cases full audio clarity was noted, 39 more were coded as most of the recording officer's audio was captured, or slightly more than 85 percent of the videos reviewed had fully or mostly discernible audio accounts of the recording officer's spoken words. If we dichotomize officer and participant audibility (1 = full/most, 0 = else) a chi-square test indicates that they are dependent χ2 (1, n = 98) = 29.4, p < .001, suggesting that as surmised, the audio channel of the BWC has superior reliability for officers.
Overall Assessment of Audio Quality for Participants and Recording Officer.
A related line of inquiry involves enumeration of the sources that led to degradation in the audio clarity. Regarding participants’ speech, any interference was coded, and interference that lasted for more than 25 percent of the entire video was also noted in Table 3. For example, distance between the participant and officer was noted as problematic in 46 of 102 videos, and it presented an issue for 25 percent or more of the video's duration in 22 instances. Background noise presented an issue in 65 videos and was an issue for 25 percent or more of the video's duration in 38 instances. When the presence of background noise is used to predict dichotomized officer audibility χ2(1, n = 98) = 10.4, p < .01, and participant audibility χ2 (1, n = 102) = 6.1, p < .01, there is significant dependence. With regard to dichotomized participant audibility, we also find significant dependence with distance χ2 (1, n = 102) = 10.5, p < .01, as well as statistical dependence with participant's speech χ2 (1, n = 102) = 4.2, p < .05. These are unlikely to be randomly distributed in police encounters. Specifically, background noise is likely a substantial problem in traffic encounters on busy streets, whereas indoor encounters in private spaces are likely to feature minimal distances or extraneous noise. To test this, we dichotomized video into those occurring “on the street” (n = 40) against those not occurring on the street (n = 62, includes indoors and outside, in yards or on porches) and ran a series of tests comparing sources of audio degradation presence to determine if the location is systematically related to problems in the audio channel. Background noise was noted as a problem in 83 percent of on-the-street videos, and the comparison with non-street videos yielded a significant contrast χ2 (1, n = 102) = 10.0, p = .002. Similarly, distance was noted as a source of audio problems in street encounters 58 percent of the time, yielding χ2 (1, n = 102) = 4.1, p = .043, indicating statistical dependence. Thus, reliability of the audio channel, that is completeness due to absence of interference, is systematically linked to location and encounter types. If we return to the differential reliability for the recording officer as compared to participants, we find the street encounters account for a significant proportion of officer audio loss χ2 (1, n = 98) = 4.1, p = .043, but the difference, although in the same direction, is not statistically dependable for participants χ2 (1, n = 102) = 3.0, p = .086.
Sources of Audio Quality Degradation for Participants and Recording Officer.
Overall Video Clarity
The quality of the video obtained from the cameras in 102 videos was assessed in two primary ways. First, lighting, color, light sources, and obstructions in all 102 videos were evaluated. This allows for a sense of how often BWC video suffers from quality issues in the sample and yields a suboptimal data source in the form of the unenhanced video. Results of the initial quality assessment are presented in Table 4. Second, the adequacy of video for describing citizen participants’ characteristics (e.g., apparent age, sex, and race/ethnicity) as judged by the coders and factors reducing this capacity were explored.
Overall Assessment of Video Quality (n = 102).
Note: aMultiple light sources are possible in a single video, hence the total sums to more than 100%.
Coders noted obstructions that concealed BWC footage in 28 percent of videos (n = 29) and obstructions that persisted for more than 25 percent of the duration of the video that occurred in 10 instances (10%). The primary sources of obstruction were noted to be officer's body parts, clothing, and equipment.
Additional examination of the video quality assessed the adequacy of video with respect to determining citizen characteristics. Citizen participants (not including officers or other service providers, but victims, suspects, and people in need of assistance that call 911) were visible in 89 videos in the sample. These videos were isolated in our analysis and VBSSO coders assessed whether the video was of sufficient quality to make a description of citizen participants (Table 5).
Video Quality for Determining Citizen Participants’ Characteristics and Sources of Interference (n = 89).
Note: aMultiple interference sources are possible in a single video, hence the total sums to more than 100%.
Coders determined that 9 percent of videos (n = 8) yielded video footage that allowed no identification of citizens from the recording. In 62 percent of cases (n = 55), the observer determined that citizen characteristics such as apparent gender, age, and ethnicity/race could be guessed (not confidently) by viewing the video. Finally, the sources of interference that contributed to the inability to confidently discern citizen characteristics were noted. Lighting (40%, n = 36) and distance (32%, n = 28) were common causes of interference with identification of citizen characteristics, however, surprisingly, camera angle, in nearly half of the videos (n = 44) was noted as an obstacle to making a confident description. If we dichotomize videos into those where confident description can be made (n = 26) compared to those of lower quality (n = 63) we find that confident descriptions are possible in 35 percent of non-street videos, but 20 percent of on-the-street videos χ2 (1, n = 89) = 2.4, p = .124, consistent with findings above with respect to direction, but not statistically significant.
Time of day (lighting) and outdoor encounters (distance) may systematically be correlated with what VBSSO can derive from the BWC data source, suggesting that triangulation of citizen characteristics, from official police records for example, would be an important consideration for researchers.
Comparison of IPSSO and VBSSO
As noted, VBSSO and IPSSO have addressed parallel questions, but have not assessed concordance and discordance across modalities. To address this gap, the 46 encounters from IPSSO and VBSSO were aligned for a comparison across commonly measured SSO elements. The IPSSO had nested elements of citizens coded within encounters allowing for coding age, race, gender, and apparent social class, and assessing each observed citizen's interaction (e.g., whether arrested, handcuffed, searched, and disrespected) with officer/police actions tied to each individual in the IPSSO. The encounter-based VBSSO did not allow for citizen-level assessments to the same degree because of difficulties ascertaining who was speaking, citizen characteristics, and the like, as noted in the previous section. Thus, the unit of analysis for VBSSO was the encounter itself (an aggregate of all citizens interacting with police at a particular event). VBSSO posed questions of whether police arrested anyone, if officers were disrespectful toward any citizen, if any citizens were in conflict, and so on. Aggregation of IPSSO citizens’ experiences up to the common level of encounter was used to better match and merge across observation conditions. Based on the results from the reliability study above we surmise that SSO elements reliant on spoken words will have higher concordance for officer and citizen comparisons and less concordance will be observed for elements reliant on the visual record.
The second issue regarding comparability stems from the IPSSO being tasked to follow the “lead decision-maker” in the two-officer team during an encounter. Due to LAPD's “contact and cover” tactical approach, this was often clearly stated when officers began a shift and if it was to be altered it was decided prior to exiting the patrol car. The video, in contrast, offers no such clear signifier to a viewer, and the researcher conducting VBSSO was not privy to the coding of the IPSSO and was never the same observer. This accommodation of aggregation is one that may yield noncomparable results when there are many citizens on-scene and some interact extensively with the “cover” officer. This will be viewable on video, but likely not a focal point for the IPSSO coder, tasked with trailing the decision maker upon exiting the vehicle. Regardless, the two sources should allow for preliminary assessment of IPSSO and VBSSO illuminating areas of strength and weakness for the latter in terms of its value as a platform for future VDA.
To assess the correspondence between VBSSO and IPSSO, a series of parallel dichotomous indicators were coded from the observation data for each of the 46 encounters. The items are consistent with the current areas being addressed by VBSSO research including indicators of police use of coercion (e.g., police verbal threats), procedural justice (e.g., police displays of respectful treatment toward citizens), and citizen states and behaviors (e.g., citizens being apparently intoxicated). These data are presented in the forms of comparison of coding (1 = yes, 0 = no, not present) whether the presence of an indicator was observed in each form of SSO. Next, a concordance measure, which computes the percentage of agreement between IPSSO and VBSSO is computed. Finally, the last two columns present Gwet's Agreement Coefficient (AC) which approximates a measure of observer agreement correcting for chance concordance and the lower bound of Gwet's AC point estimate within the 95% confidence interval (Gwet 2014, 2021; Klein 2018). The comparison measures include three domains where both modalities of SSO have played a role including police legal and coercive actions, procedural justice, and citizen states and behaviors.
The IPSSO and VBSSO comparison commences with legal and coercive actions presented in the top panel of Table 6. Results indicate that, in reference to Landis and Koch's (1977) criteria for benchmarking agreement, the paired observations have substantial to almost perfect agreement regarding whether an encounter involved an arrest, handcuffing, reaching for or brandishing a weapon, whether police issued a threat to a citizen, or had a firm grip hold on a citizen. Finally, whether the police conducted a search had only fair agreement across the SSO modalities, however, IPSSO observation yielded nearly twice as many encounters (n = 17) as VBSSO (n = 9) where a search was determined to have been conducted. The VBSSO, with multiple records in two-officer units, does offer a comprehensive view of police interaction on a scene, rather than the IPSSO which is limited to a primary focus on the single officer. Nevertheless, a general pattern emerges, that IPSSO generally yields more instances of the presence of each indicator in the encounters in the sample. We cannot discern whether this is definitively overcounting IPSSO or undercounting VBSSO, but it is suggestive of a divergence in the methods.
Legal/Coercive Actions, Procedural Justice, and Citizen States in Police-Citizen Encounters (n = 46).
Procedural justice (coding exemplars can be found in Appendix 2) is captured in four measures in the middle panel of Table 6. The concordance of police disrespect is high, and Gwet's AC indicates substantial agreement on that measure. Participation, neutrality, and indicators of police respect, however, only generate slight to fair levels of agreement as measured by the lower bounds of Gwet's AC. This is at odds with our supposition that measures rooted largely in the spoken word would show higher concordance and suggests that subjectivity may also be higher in these indicators.
The last domain of measures, citizen states, and behaviors, includes six indicators of citizen state, presented in the bottom panel of Table 6. These were coded with regard to whether any citizen present showed evidence of injury or illness, apparent mental illness, influence of alcohol or drugs, conflict with police, disrespect toward police, or respectful behavior toward police. In comparisons of citizen states across SSO conditions, almost perfect agreement is found in identifying the presence of injured or physically ill citizens in encounters and regarding evidence of drug or alcohol use. In the latter, however, VBSSO yielded no cases where citizens had apparent alcohol or drug use compared to three in IPSSO, offering some evidence that the methods diverge in this area. Moderate levels of agreement were obtained for whether a citizen was observed to be mentally ill, disrespectful, or in conflict with the police. Finally, citizen respect showed a fair level of agreement, according to the statistical benchmark, and was one measure, along with police disrespect, where VBSSO produced a greater number of observed indicators when compared with IPSSO.
The single-item comparisons of concordance suggest that the complexity of encounters and the quality of BWC audio/video may contribute to discrepancies between the SSO modalities. Though the sample is small (n = 46), indices of concordance were calculated for police coercive actions (summing seven binary concordance scores for handcuffing, displaying weapon, and other action.; x̅ = 6.3), four indicators of procedural justice concordance (x̅ = 3.07), and six indicators of citizen states and behavior concordance (x̅ = 5.04) to generate three dependent measures in Table 7. These measures were then regressed on indicators of BWC quality and the complexity of the coded encounter (results are not shown, Supplemental materials outline measures and results are shown in Appendix 3). Two measures of BWC quality focus on contrasting those cases with 25 percent audio interference (17% of cases) and if any video was obstructed for 10 percent or more of the encounter (28% of cases). Four measures of complexity reflect the count of citizens (x̅ = 2.98) and police on scene (x̅ = 3.26), binary measures of whether the encounter occurred on the street (39% of cases), and whether the encounter lasted less than 15 min (35% of cases) round out this domain representing the complexity and possibly attention fatigue for observers.
Descriptive Statistics: Concordance Measures, Audio/Visual Problems, and Complexity (n = 46).
Note: IPSSO = pioneered in-person systematic social observation; VBSSO = video-based systematic social observation.
Ordinary least squares (OLS) regression models were estimated on all three concordance indices, with resulting dot-and-whisker plots presented for models using point estimates and unadjusted standard deviations as shown in Figure 1. Bonferroni adjustment of p-values yielded only one statistically reliable estimate, with the number of police on scene having a substantially negative impact on concordance for police coercion. Given the small sample size, patterns of coefficient signs offer a second evaluative approach. Specifically, all six audio/video problem estimates indicate reduced concordance, which is a pattern consistent with expectations. Shorter encounters, those less than 15 min suggest positive concordance, consistent with lower complexity, and the pattern of street encounter coefficients suggest the anticipated attenuation of concordance.

Estimates and unadjusted standard deviation dot-and-whisker plots for ordinary least squares (OLS) coefficients.
An alternative approach to comparing the VBSSO and IPSSO concordance is to consider that individual items vary along the dimensions of high subjectivity and low subjectivity. Put differently, some items require applying a more complex set of decision rules to observed behavior and might be considered latent items, compared to explicit or manifest measures with lower subjectivity. To address this difference (which is embedded in items in various indices), we examined the content of definitions for indicators. To account for subjectivity, we opted for six-item scales reflecting higher or lower subjectivity as shown in Table 7. Lower subjective items include: arrest, handcuffing, brandish weapon, reach for weapon, any citizen injury, and whether a search occurred. Higher subjective items include citizen displays of respect and disrespect, police displays of respect and disrespect, whether citizens exhibited signs of mental illness, and whether any citizens displayed conflict toward the police. The pattern of findings, reflected in the lower two plots as shown in Figure 1, suggests a similar pattern to the regressions noted above. However, there is no statistically significant predictor of either index once Bonferroni correction adjustments were applied to significance levels.
Discussion and Research Implications
This article examined the challenges, strengths, and the limitations of BWC footage as a source of data for future research. We examined issues of BWC validity in the context of VDA by examining audio and video problems and its reliability by directly comparing the coding of footage from a sample of police-citizen encounters to coding of the same events by on-scene coders.
We found that there are substantial and systematic audio and video gaps present in BWCs as a source of data for social science investigation. These gaps are likely to impact the reliability and validity of measures developed from those data. Problems emerge because of the way a camera and its microphone are positioned on an officer. The video yields differential incompleteness, or biased and imbalanced data source, regarding the record of participants’ verbal statements (63% of videos feature most or full audio of participants’ utterances) as compared to the recording officer (86% feature most or full audio of the recording officer). The video capture imbalance was recognized by Makin et al. (2019), for example, in the coding of emotion states of citizens and police. Noting the findings of Aviezer, Trope, and Todorov (2012) regarding the importance of body cues to the interpretation of emotions, they posited that a lack or limited number of such cues for police, since they wear the camera, is an important limitation for researchers to consider. Similarly, our analyses suggest BWC is prone to missing data, as evidenced from comparison coding of visible police legal and coercive actions; it appears to undercount those police actions consistent with the concern of suboptimal capture. However, our analysis suggests BWC has substantial reliability in capturing police utterances.
One research project that would make sense for VBSSO researchers to undertake is the comparison of single and multiple camera coding from the same event. That is, what does a second camera add or enhance in terms of the measurement of aspects of the encounter? Our analysis here has been reliant on two-officer BWC due to the LAPD deployment approach. However, an important reliability question can be answered by identifying lead officers in a sample of events and conducting a project coding that single video and then comparing those results with VBSSO coding using the same items with the second officer's camera. This would illuminate the gains of using a second video, and the shortcomings of a single video account, and might also identify measures where the cost of coding a second video would not yield significant improvements.
The reliability of BWC as a timestamped record for VBSSO, however, should not be understated. Having accurate dates and times within the audio and video provide researchers with substantial capacity for judging sequential developments, causal ordering, and the duration of events (Schafer, Hibdon and Kyle 2022; Willits and Makin 2018) in the record preserved in BWC data. Thus, the technology should open theoretical frames that are more efficient and less cumbersome than in-person coding. Voigt et al. (2017), for example, perform sequential utterance level analysis, which is supported by the results here. Further, the ability to estimate inter-rater reliability in a VBSSO setting represents an advance that has been absent in IPSSO and will, as projects share results, yield greater consistency and comparability of measurement across projects and over time. This may also lead to the reduction of subjectivity in measurement which is difficult to judge in the IPSSO setting.
Advances in computing are quickly allowing for the possibility of computer-aided VDA which can be applied to VBSSO using BWC footage. Automation of audio and video coding of BWC video are at different stages given the limits of technology and the quality of the data collected. Natural language processing (NLP), for example, allows for the audio track to be converted to a text stream and unleashes a variety of readily available search possibilities within the transcript of a video (e.g., slurs or profanity, or keywords such as gun, weapon, knife, taser, and punch). These can be used to classify videos for further viewing and perhaps to help identify specific timestamps of segments for review. NLP is likely to have the capacity to generate new approaches to analysis of the audio track that have been suggested by Voigt et al.'s (2017) research. The second area of computer-aided advancement is in computer vision (CV). At this writing CV has shown some capacity for recognition of images and actions with growing reliability, however, the BWC presents a compound problem of not necessarily observing the action due to mounting location or obstruction and image jitter due to the motion of the wearer. Currently, these issues make automated classification of BWC footage more computationally difficult than footage captured from fixed cameras more typically used in CV applications. Indeed, VDA researchers in other areas are successfully deploying these techniques (Nassauer and Legewie 2022, Ch. 8). As NLP and CV advance computationally, we anticipate both to become powerful tools for the identification and classification of BWC videos for review and analysis (e.g., those with conflict, escalation, and de-escalation) and to open new possibilities for analysis in the coming years.
We suggest that the theoretical development and testing of VDA applications are important pathways for future researchers (see also, Makin, Willits and Brooks 2020). Nonetheless, other data sources should be used to assess validity and reliability of encounters or events. In police agencies, calls for service, incident reports, traffic stops, and field interviews likely have data on race, sex, and age of participants in many encounters, and could provide important checks on coding. Triangulating these data with BWC footage will provide increased validity.
We also strongly suggest that researchers report on a number of key areas related to sampling and completeness of encounters. For example, researchers should report on how videos were sampled; that is, were they sampled from the metadata on platforms created by the vendors? Was the sample based on department-wide incidents or by specific patrol divisions or by special units? Further, it is important to know the number of videos available per event, the attrition levels of videos/events, an assessment of audio/video problems, and the level of confidence that coders have in the assessments being made regarding variables. Researchers should explain how BWCs are used within policing agencies, based in part on departmental policies (e.g., when they are activated, percentages of compliance, or if citizens are notified that they are being videotaped). Researchers should fully disclose the challenges and limitations of the footage that they acquire.
From this study and from our work on IPSSOs, we believe that using VBSSOs of BWC may have significant cost savings over IPSSO. This is particularly the case in rare events. To our knowledge, no officer-involved shootings have been captured in IPSSO and this represents one of the many areas where this research suggests the utility of thoughtful VDA application.
Within this article, we focused on the use case of BWC for VDA, with an emphasis on police agencies in the United States. The extent to which this generalizes to the use and access in other locales varies and is dynamic. Access, release, and retention even within police departments, over time, is likely to change. Other use cases for BWC have not been studied or mentioned here, but we would argue that the findings are relevant to other areas. In particular, the general view that verbal records will be of higher quality than the video will be like other use cases (e.g., nursing, business meetings, or teaching environments). More important than the extension of these findings to other use cases is the recognition of the ubiquity of practice in U.S. policing making reactivity to the camera a limited concern for the wearer or those being filmed. This is unlikely to be true for novel uses where the presence of the camera may elicit performative responses.
The second area is the one that concerns the ethics and consent of those captured by novel BWCs. Viewing, coding, and storage of BWC footage is currently under extensive legal scrutiny regarding the balance of privacy and public disclosure. At the same time, however, state legislatures are now mandating the collection of footage throughout the United States. In projects where researchers exercise direct control of the video, storage, and access, the concern for personal revelations, ethics of disclosure, and the possibilities of research-based harm must be carefully considered and mitigated. The oversight by institutional review boards for researchers becomes an important method of accountability along the line of cost–benefit assessments involving BWC and VDA.
Supplemental Material
sj-docx-1-smr-10.1177_00491241231156968 - Supplemental material for Video Data Analysis and Police Body-Worn Camera Footage
Supplemental material, sj-docx-1-smr-10.1177_00491241231156968 for Video Data Analysis and Police Body-Worn Camera Footage by John D. McCluskey and Craig D. Uchida in Sociological Methods & Research
Footnotes
Acknowledgments
The authors would like to thank the anonymous reviewers who provided us with helpful comments and suggestions that improved our original manuscript and its contribution to the field. There are many people who played important roles in the research that was undertaken and we thank them for assisting us. In particular, we thank William Lu, Crime and Intelligence Analyst with the LAPD, and JSS Research Associates Mariel Shutinya and Lauren Revier for their diligence in coding and reviewing footage. Special thanks go to NIJ's former Director Nancy Rodriguez and Senior Computer Scientist Joel Hunt. Within the LAPD, we are grateful to former Chief Charlie Beck and current Chief Michel Moore, Deputy Chief Sean Malinowski (ret.), CIO Maggie Goodrich (ret.), Lt. Daniel Gomez (ret.), Deputy Chiefs Jorge Rodriguez and Michael Rimkunas, and Commanders Todd Chamberlain (ret.) and Robert Marino for their support in allowing us to ride with and observe officers. Last, but not least, we thank the patrol officers and supervisors of Newton and Mission Divisions for accepting us into their vehicles and lives (however briefly) to ride with them and begin to understand their work.
Appendices
Data Availability
Data are archived in the following figshare locations with explanations and command files for execution:
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This project was supported in part by Award No. 2014-R2-CX-0101 by the National Institute of Justice, Office of Justice Programs, U.S. Department of Justice to the Los Angeles Police Foundation, and Justice & Security Strategies, Inc. This article was also supported in part by the Bureau of Justice Assistance (BJA), CNA Corporation with Justice & Security Strategies, Inc. (JSS), and Arizona State University (ASU) as sub-recipients (grant number 2019-BC-BX-K001). BJA is a component of the U.S. Department of Justice Office of Justice Programs. John McCluskey's work on this paper was supported, in part, by the RIT Paul A. and Francena L. Miller Research Fellowship. Points of view or opinions contained herein do not necessarily represent the official position or policies of the U.S. Department of Justice.
Supplemental Material
Supplemental material for this article is available online.
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
