Abstract
How online social behavior covaries with real-world outcomes remains poorly understood. We examined the relationship between the frequency of misogynistic attitudes expressed on Twitter and incidents of domestic and family violence that were reported to the Federal Bureau of Investigation. We tracked misogynistic tweets in more than 400 areas across 47 American states from 2013 to 2014. Correlation and regression analyses found that misogynistic tweets were related to domestic- and family-violence incidents in those areas. A cross-lagged model showed that misogynistic tweets positively predicted domestic and family violence 1 year later; however, this effect was small. Results were robust to several known predictors of domestic violence. Our findings identify geolocated online misogyny as co-occurring with domestic and family violence. Because the longitudinal relationship between misogynistic tweets and domestic and family violence was small and conducted at the societal level, more research with multilevel data might be useful in the prediction of future violence.
Approximately 30% of women worldwide reported experiencing abuse at the hands of an intimate partner (World Health Organization, 2017). In the United States, the Federal Bureau of Investigation (FBI) collects data on arrests for offenses against family and children, which are “unlawful nonviolent acts by a family member (or legal guardian) that threaten the physical, mental, or economic well-being or morals of another family member and that are not classifiable as other offenses, such as Assault or Sex Offenses” (U.S. Department of Justice, 2013, p. 165). When these forms of aggression are directed toward romantic partners, they are referred to as intimate-partner violence (Adams & Beeble, 2019; Ali, Dhingra, & McGarry, 2016). Most of these arrests are made for violence toward women. Approximately 70% of all victims of family violence are women, and approximately 70% of offenders are men (Durose et al., 2005; FBI, 2020). Known risk factors for domestic violence include gender inequality, income inequality, alcohol availability, low education, exposure to domestic violence, childhood abuse, ethnicity, and place of residence within the United States. Krahé (2018) proposed an organizing framework for understanding these risk factors, dividing them into the individual, intimate relationship, and societal levels. In the current research, we investigated a novel societal-level risk factor. Specifically, we examined the extent to which misogynistic posts on a widely used social networking platform correlate with and predict domestic and family violence in the United States.
The American Psychological Association (2020) defines misogyny as “hatred or contempt for women.” Misogynistic norms typically manifest and are reified through disrespectful, hostile, sexist sentiments and attitudes (e.g., sexually objectifying women, endorsing male dominance over women, endorsement of rape culture; Leone & Parrott, 2019). Similarly, misogyny is positively associated with sexual and physical aggression toward women (Malamuth, Linz, Heavey, Barnes, & Acker, 1995; Parrott & Zeichner, 2003). Thus, many men who hold hostile or misogynistic views toward women are likely to endorse and engage in violence toward them.
Misogyny and domestic and family violence are global phenomena. Studies in many societies show that attitudes that are accepting of violence and gender inequality are correlated with higher levels of violence toward women. A meta-analysis of 85 studies found a large effect of attitudes condoning marital violence on perpetration of domestic violence and a moderate effect for traditional sex-role ideology (Stith, Smith, Penn, Ward, & Tritt, 2004). In one large-scale study of more than 20,000 people in eight African nations, misogynistic attitudes regarding women’s sexuality (e.g., women do not have the right to refuse sex) and attitudes condoning violence toward women (e.g., women deserve to be beaten) also correlated with domestic violence (Andersson, Ho-Foster, Mitchell, Scheepers, & Goldstein, 2007). These effects were robust to several known causes, including age, education, household income, and food security. Correlations between attitudes accepting of violence and hostile attitudes toward women are not limited to developing nations. These relationships have been observed in large samples of North American college students (Reitzel-Jaffe & Wolfe, 2001) and in a study with more than 70,000 respondents from 51 nations (Herrero, Rodríguez, & Torres, 2017). In sum, misogynistic and violent social norms may be significant contributors to violence against women.
We suspect that social media platforms are a powerful means of contagion of misogynistic attitudes and spiteful feelings, even more so than face-to-face interactions. The phenomenon of emotional contagion is the tendency to automatically mimic verbal and nonverbal cues and synchronize expressions in a way that creates emotional convergence with another person (Hatfield, Cacioppo, & Rapson, 1993). Goldenberg and Gross (2020) noted that although emotional contagion via social media is quite similar to face-to-face contagion, there are some important differences between the two phenomena. Social media provide many more opportunities for exposure to misogynistic content than face-to-face interactions and the ability to reach more people. Some social media users are motivated to express strong emotions and attitudes, a strategy that attracts attention on a platform in which users compete for attention (Goldenberg & Gross, 2020). Likewise, misogynistic attitudes are characterized by anger (Parrott & Zeichner, 2003) and aggressive humor (Siebler, Sabelus, & Bohner, 2008), and anger spreads faster than joy among social media users (Fan, Xu, & Zhao, 2016).
Statement of Relevance
Violence is a significant social issue and source of concern to the general public as well as the subject of research in many scientific disciplines. In this research, we investigated whether violence directed toward the family, as reported to the Federal Bureau of Investigation, was related to misogynistic messages in social media posts (e.g., tweets). Because more than 70% of reported victims of family violence are women, this study focused on the relation between misogynistic messages and violence against women in particular. We found that the strongest predictor of family violence at Time 2 was family violence at Time 1, 1 year earlier. Misogynistic messages at Time 1 were also a significant predictor of family violence 1 year later at Time 2, although the effect was small. These findings do not indicate that misogynistic messages on social media cause domestic violence. Yet they do suggest that the expression of prejudicial messages against women systematically co-occurs with domestic violence.
We propose that misogynistic attitudes in the online world might spill over and incite some men to off-line domestic and family violence. Alternatively, clusters of real-life violence may show up in online spaces after the violence occurs. Some research shows, however, that exposure to social media does influence online and off-line behavior. For instance, tweets and Facebook messages aimed at effecting behavioral change influenced real-world health behavior (Centola, 2010) and voting in the 2010 U.S. congressional elections (Bond et al., 2012). Bond et al. manipulated political-mobilization messages on Facebook and found that the greatest changes in political behavior were among users and their close friends. Most contagion takes place within close-knit communities because these individuals tend to be similar to each other, and tweeting and retweeting facilitates social bonding (Harrigan, Achananuparp, & Lim, 2012). Anger toward women and aggressive misogynistic humor on social media also likely facilitate bonding among men (Thomae & Pina, 2015). Thus, social media may maintain the salience of misogynistic social norms, which may incite some men to off-line violence toward women.
If misogynistic attitudes lead to more violence toward women, such violence should occur more frequently in geographic areas where people share more misogynistic online content. In support of this notion, one study found that misogynistic tweets positively correlated with FBI statistics on sexual violence across the United States (Fulper et al., 2014). Although this study provided a promising initial test of the relationship between social media posting and real-world violence, it suffered from a number of limitations, including data broadly aggregated to the state level, which prevented a high-powered test of online misogyny and violence. The researchers likewise neglected important societal-level predictors of violence and operationalized misogyny using terms that were potentially ambiguous and unvalidated. Further, the researchers examined correlations between misogyny online and sexual violence using just 1 year of cross-sectional data, which prevented an examination of causal relationships over time. They also did not examine the FBI crime statistics for domestic and family violence (which is more frequently reported than sexual violence). The current research addressed these limitations.
The Current Research
We examined the relationships between misogynistic tweets and domestic and family violence in the United States during 2013 and 2014. To maximize power and generalizability, we used relatively small core-based statistical areas (CBSAs), which provide good geographic granularity. Our research is also unique in that it longitudinally tested the possibility that prior exposure to misogynistic tweets would correlate with future domestic and family violence using a cross-lagged model. We also controlled for several variables known to influence this type of violence. We predicted that the number of misogynistic tweets would positively correlate both cross-sectionally and longitudinally with domestic- and family-violence incidents.
Method
We combined three elements to conduct our analyses: misogynistic tweets, FBI crime reports, and covariate data from the American Community Survey, which contains data on sociostructural characteristics. The term CBSA refers collectively to micropolitan statistical areas and metropolitan statistical areas, which are defined by the U.S. Office of Management and Budget. CBSAs are determined by population. Micropolitan statistical areas are urban clusters of at least 10,000 but fewer than 50,000 people. Metropolitan statistical areas are urban clusters of at least 50,000 people. Currently, there are 917 defined CBSAs in the United States. The sample size was determined by the number of CBSAs for which all data were available during 2013 and 2014 (i.e., crime statistics, misogynistic tweets, and covariates). We used this date range because it was the only available range for which all of the requisite covariates and crime data were available. The number of misogynistic tweets, crime data, and sociostructural measures were all at the CBSA level. We reasoned that 2 years’ worth of tweets from more than 400 CBSAs was sufficient to provide an indication of online misogyny in each geographic area. All measures collected are reported here, and all data and R code can be found at https://osf.io/sgqav/.
Twitter data
Misogyny-term generation
Using hashtag finders and manual Twitter searches, we generated a list of hashtags and sentences that contained misogynistic sentiment. Hashtag finders create networks of hashtags to identify those hashtags commonly posted together on social media platforms, and our manual searches achieved the same aim on Twitter. We also compiled a list of misogynistic and condescending terms for women. We then combined the sentences with the condescending female terms to generate a list of search terms that were clearly misogynistic (e.g., “I hate women,” “bitch get back in the kitchen,” “you asked for it skank,” and “make me a sandwich slut”). We chose to combine some terms into a short sentence to reduce ambiguity and increase validity. Frequency of the search terms appears in Table 1, and a list of all terms and term combinations is in Table 2. We conducted a manual validity check of 1,000 randomly selected tweets to confirm misogynistic content. To do so, two blind coders read each tweet and judged whether it expressed misogyny. Both raters judged 92.7% of tweets as unambiguously misogynistic. Given the necessary trade-off between sensitivity and specificity, we decided that this error rate and its accompanying noise were sufficient for the analyses we performed.
Frequency of Misogynistic Search Terms
All Terms and Combinations of Misogynistic Search Terms
Application of terms to Twitter corpus
We applied these search terms to a public corpus of 1.8 billion tweets downloaded via the Twitter application programming interface between January 2013 and December 2014, inclusive. This corpus contained tweets from the “sprinkler stream,” which is free to the public and contains a random collection of approximately 1% of the full Twitter stream. Restricting this stream to tweets containing only our highly specific list of misogynistic sentences resulted in 16,791 tweets across the 2 years.
Geolocation of tweets
We geolocated these tweets to the U.S. city level using a validated Twitter geolocation algorithm (Blake, Bastian, Denson, Grosjean, & Brooks, 2018). The algorithm compared user-defined strings in the user’s location field with a dictionary of 5,567 U.S. cities, villages, boroughs, towns, and census-designated places (all locations in the United States with populations reported to exceed 5,000 people). Of the 16,791 tweets, 3,472 were geolocated to the U.S. city level (1,538 cities). These city-level tweets were then aggregated to their respective CBSA so the tweet data could align with the FBI crime data (which were available only at the CBSA level). Ninety-six percent of the city-level tweets (3,317 tweets) were matched to a CBSA.
Covariates
We controlled for several variables that often covary with domestic violence. These variables included U.S. region, structural inequalities, alcohol availability, race, education, income, population size, and age. These sociostructural variables were collected at the CBSA level from the 2013 and 2014 U.S. Census Bureau 1-year estimates from the American Community Survey (U.S. Census Bureau, 2017). We selected the covariates because domestic violence varies as a function of age (Peters, Shackelford, & Buss, 2002), race (Smith et al., 2017), income and education (World Health Organization, 2017), gender inequality (Archer, 2006), income inequality (Yapp & Pickett, 2019), alcohol (McKinney, Caetano, Harris, & Ebama, 2009), and population. The household Gini coefficient was our measure of income inequality. For educational attainment, we measured the proportion of the population that had some college education. The income variable measured the median earnings for the civilian population age 16 years and older. We also obtained median age, the percentage of the population that was White, and the total population of individuals age 16 years and older. The alcohol variable was the number of beer, wine, and liquor stores specifically designated for off-premises sales.
We operationalized gender inequity using four measures and according to the following dimensions from the United Nations Gender Inequality Index. For the reproductive-health dimension, we measured the percentage of people without health insurance who were women. For the empowerment dimension, we measured the percentage of people who had achieved some college education who were men. For the labor dimension, we measured the percentage of the combined nonfamily household male and female income attributed to men and the percentage of unemployed people ages 20 to 64 years who were women. For all inequality measures, higher scores reflected more gender inequality.
FBI crime data
We gathered data from the FBI Uniform Crime Reporting Program (FBI, 2015, 2016). The FBI reporting system does not separate intimate-partner violence from other forms of domestic and family violence. We therefore relied on the data for FBI Offense Code “Offenses Against the Family and Children” for the years 2013 and 2014. These crime data were aggregated to the CBSA level, across months, to form yearly estimates.
Statistical analyses
All data were analyzed in the R programming environment (Version 3.6.1; R Core Team, 2019). We obtained data from CBSAs in 47 states. All data were gathered for the years 2013 and 2014. The Twitter data and the data on domestic-violence incidents were collected throughout each month of these years and then aggregated to the year, whereas the covariate data were available only as annual summaries. Cross-sectional analyses combined both years and included zero-order correlations and regression of violence on the covariates and Twitter data. We also tested a cross-lagged model, in which 2013 tweets (Time 1) and domestic-violence incidents longitudinally predicted these outcomes in 2014 (Time 2). The data sets for Time 1 and Time 2 contained 441 and 454 CBSAs, respectively. We also created dummy variables for U.S. region determined by the U.S. Census Bureau, which divides the country into four regions: Northeast, South, West, and Midwest. We specified the Midwest as the reference region. The number of alcohol outlets and population were natural-log-transformed because of skewness. We computed descriptive statistics and zero-order correlations among the variables. All variables except domestic violence were z-transformed prior to being entered into regression analyses.
Because the FBI reports crime data as the number of incidents in each CBSA, we used negative binomial regression, which is appropriate for modeling count data. We used the glm.nb function in the R package lme4 for the regression analyses (Bates, Mächler, Bolker, & Walker, 2015). Our primary analysis consisted of a negative binomial model predicting domestic and family violence from the number of misogynistic tweets and covariates related to violence and women’s status. Across the CBSAs, there were only 33 and 35 instances of zero counts for domestic and family violence at Time 1 and Time 2, respectively (i.e., less than 7% of CBSAs); thus, our data were not zero-inflated. Nonetheless, we examined the comparative fit of our negative binomial model with zero-inflated and Poisson models. This analysis showed that the Akaike information criterion (AIC) was much larger for the zero-inflated model (AIC = 68,618) than the negative binomial model (AIC = 8,343). Vuong’s (1989) test showed that the negative binomial model had a significantly better fit than the zero-inflated model (z = 13.67, p < .0001). Thus, the zero-inflated model was inappropriate for our data. The Poisson model revealed significant overdispersion (z = 9.38, p < .0001), which made it inappropriate for our analyses. Moreover, the AIC was much larger for the Poisson model (AIC = 73,804) than our negative binomial model, demonstrating an inferior fit.
We followed the regression analysis with a cross-lagged panel analysis to determine whether domestic violence and misogynistic tweets at Time 1 would predict these outcomes at Time 2. We used the lavaan structural equation modeling package in R (Rosseel, 2012). Because the path modeling could not simultaneously include maximum likelihood and negative binomial regression, we standardized all variables. However, to account for the nonnormality of the count data, we used Huber-White robust estimation of standard errors (White, 1982). Covariates of each 2014 outcome included the 2013 values for population; alcohol outlets; percentage of White residents; gender inequality in education, employment, health, and income; the Gini coefficient; percentage of the population who were college educated; median age; and median income.
Results
Descriptive statistics and correlations
Table 3 shows the means, medians, standard errors, minimum values, and maximum values for all the variables. Table 4 displays the zero-order correlations among the variables. Domestic violence was positively correlated with the number of misogynistic tweets, r(893) = .51, p < .0001; alcohol outlets, r(890) = .52, p < .0001; population, r(858) = .20, p < .0001; percentage of people with some college education, r(893) = .21, p < .0001; percentage of White residents, r(861) = −.14, p < .0001; gender inequality in education, r(893) = .08, p = .016; and median income, r(893) = .10, p = .003. The strongest correlates of misogynistic tweets were the number of alcohol outlets, r(890) = .55, p < .0001, and inversely, percentage of White residents, r(861) = −.23, p < .0001. Despite significant zero-order correlations among the predictors, all variance-inflation factors in subsequent analyses were below 2.06.
Descriptive Statistics for Study Variables
Note: Population and alcohol outlets are presented untransformed for ease of interpretation. GI = gender inequality.
Correlations Between Study Variables
Note: GI = gender inequality.
p < .05. **p < .01.
Primary analysis
The counts of domestic-violence cases were overdispersed, as indicated by the variance (21,213.57) of this variable being much greater than the mean (79.92). Thus, a negative binomial model was appropriate. Sixty-four cases were omitted because of missing data, leaving 831 valid cases. Table 5 presents the results from the full negative binomial regression model. The number of misogynistic tweets and number of alcohol outlets significantly predicted domestic-violence incidents. In terms of geographic region, domestic violence was less common in the Northeast relative to the Midwest. No other predictors were significant.
Results From Negative Binomial Regression Analyses Predicting Number of Domestic-Violence Occurrences in the United States During 2013 and 2014
Note: Population and alcohol outlets were natural-log-transformed. All variables except the outcome variable (i.e., number of domestic-violence occurrences) were standardized. CI = confidence interval; GI = gender inequality.
Path analysis
Figure 1 displays the cross-lagged panel model, which included all covariates. Time 1 domestic-violence incidents predicted Time 2 domestic violence (β = 0.96, z = 49.28, p < .001). Similarly, Time 1 misogynistic tweets predicted Time 2 misogynistic tweets (β = 0.62, z = 19.85, p < .001). In terms of the crossed paths, Time 1 misogynistic tweets predicted Time 2 domestic violence (β = 0.03, z = 2.32, p = .021); however, Time 1 domestic violence did not predict Time 2 misogynistic tweets (β = −0.02, z = −0.72, p = .47). The predictors (including covariates) accounted for 96.7% of the variance in Time 2 domestic violence and 94.5% of the variance in Time 2 misogynistic tweets.

Results of cross-lagged analyses showing the relationships between misogynistic tweets and domestic violence across 1 year. All variables were standardized. Covariates of each Time 2 outcome included the Time 1 values for population, alcohol outlets, percentage of White residents, the Gini coefficient, percentage of the population that was college educated, median age, median income, U.S. region, and gender inequality in education, employment, health, and income. Asterisks indicate significant paths (*p < .05, ***p < .0001).
We also conducted a model comparison to determine whether misogynistic tweets at Time 1 made a significant and unique contribution to explaining the variance in the number of domestic- and family-violence incidents at Time 2. Specifically, we compared fit between a model with the tweets and a model without the tweets. Both models contained all of the Time 1 covariates and Time 1 domestic- and family-violence incidents. The model with the tweets fitted the data significantly better than the model without the tweets, χ2(1) = 4.33, p = .037. This finding suggests that misogynistic tweets accounted for a significant portion of the variance in domestic- and family-violence incidents beyond Time 1 domestic and family violence and the remaining covariates.
Discussion
The present research found that domestic- and family-violence incidents across the United States were positively correlated with the number of misogynistic tweets in those areas, even when analyses controlled for several known societal-level correlates of domestic violence. Because we were able to obtain data from 47 of the 50 American states, our findings are likely to generalize to most of the United States. However, it remains to be seen whether the relationship between misogynistic tweets and domestic and family violence generalizes to other nations and other social media platforms. Furthermore, a cross-lagged regression model found that the number of misogynistic tweets in 1 year preceded and positively predicted domestic and family violence the following year. Although this latter effect was small, it remained a significant predictor of domestic and family violence in our analyses. Indeed, because domestic violence and family violence are underreported to law enforcement, our data may actually underestimate the predictive ability of misogynistic tweets on violence toward women and children.
Alcohol has a well-documented relationship with domestic violence. Our study replicated that work because the number of outlets licensed to sell off-premises alcohol was the strongest predictor of domestic and family violence. This finding is consistent with those of many other studies (McKinney et al., 2009). Presumably, this relationship is due to intoxicated individuals behaving aggressively. However, this notion is speculative because our data do not allow individual-level conclusions to be drawn. We also found positive correlations between misogynistic tweets and the percentage of college-educated people exposed to domestic and family violence, which no longer existed in the regression. Additionally, we found small negative correlations between the percentage of White people and (a) the number of domestic-violence incidents and (b) misogynistic tweets. However, when we included our covariates in the model, the coefficient for the percentage of White people, which showed only a small zero-order correlation to begin with, was no longer significant. This finding suggests that misogynistic tweets and alcohol were better determinants of domestic and family violence than ethnic composition of the CBSAs.
Our research joins a broader effort to understand the extent to which social media can be used to learn more about criminal offending and intergroup violence. Analysis of affective and moral sentiments on Twitter successfully predicted thefts in Chicago (Chen, Cho, & Jang, 2015); public disorder during events organized by a right-wing, anti-Islam group (Jurek, Bi, & Mulvenna, 2014); and violence and arrests during the 2015 Baltimore protests (Mooijman, Hoover, Lin, Ji, & Dehghani, 2018). This latter effect was greater for people who believed that their values were widely shared, thus highlighting the role of perceived social norms in Twitter use and violence perpetration. Similarly, a laboratory experiment manipulated perceptions of social norms about rape. Participants were told that acceptance of rape myths in their peer group was either high or low (Bohner, Siebler, & Schmelcher, 2006). The men who were told that rape acceptance was widespread exhibited a higher proclivity for rape. Spreading misogynistic tweets may create or maintain communities of men with shared misogynistic beliefs.
In combination with other robust societal-level predictors and further research, geolocating observed online misogyny may eventually be one factor (among many) used to inform crime-forecasting tools and interventions. To more precisely determine the roles of social media in shaping the extent and location of future violence, more research is needed with greater temporal resolution of tweets and violence (e.g., monthly) over a longer period of time (e.g., monthly measures over many years). Because the longitudinal relationship between misogynistic tweets and domestic and family violence was small, and our data were conducted at the societal level, more research with corroborating individual-level data might be useful in the prediction of future acts of violence. For instance, simultaneously obtaining data from Krahé’s (2018) three levels of analysis, such as individual-level risk factors, relationship-level risk factors (such as problematic interactions between partners), and societal-level factors (such as misogynistic tweets), would likely enhance our understanding of the relationship between misogynistic social media and domestic and family violence.
One limitation is that the FBI categorizes intimate-partner violence together with less common forms of family violence, such as child abuse and elder abuse. Although this limitation may be rectified by future work, research using FBI data found that the most common victim in family-violence offenses was the spouse, and 73% of victims were women (Durose et al., 2005). A second limitation of the data set is that the most severe acts of intimate-partner violence were not included in the analyses because the FBI excludes partner-related homicides, assaults, and sexual assaults from their tally of offenses against the family and children. However, the severity of the offense was still deemed sufficient for arrest, and severe forms of sexual and physical intimate-partner violence are highly correlated with psychological aggression (e.g., controlling behavior; Krebs, Breiding, Browne, & Warner, 2011). Thus, noise from including violence committed against nonpartner family members and excluding severe acts of violence may have underestimated the relationship between misogynistic tweets and intimate-partner violence.
Conclusion
We found that the frequency of misogynistic posts in the United States on Twitter covaried with and positively predicted the occurrence of domestic and family violence over 1 year. Our study is the first to use big data to predict domestic violence from misogynistic tweets across a 2-year period, and with it, we demonstrate the possible benefits of monitoring and geolocating social media content to prospectively predict and counter real-world violence.
Footnotes
Transparency
Action Editor: Kate Ratliff
Editor: Patricia J. Bauer
Author Contributions
K. R. Blake designed the study, obtained the Twitter data, conducted the statistical analyses, and cowrote the manuscript. J. Lian obtained the American Community Survey data and prepared the data sets for analysis. S. M. O’Dean cowrote the manuscript. T. F. Denson designed the study, conducted the statistical analyses, and cowrote the manuscript. All the authors approved the final manuscript for submission.
