Abstract
Twitter data are widely used in the social sciences. The Twitter Application Programming Interface (API) allows researchers to build large databases of user activity efficiently. Despite the potential of Twitter as a data source, less attention has been paid to issues of sampling, and in particular, the implications of different sampling strategies on overall data quality. This research proposes a set of conceptual distinctions between four types of populations that emerge when analyzing Twitter data and suggests sampling strategies that facilitate more comprehensive data collection from the Twitter API. Using three applications drawn from large databases of Twitter activity, this research also compares the results from the proposed sampling strategies, which provide defensible representations of the population of activity, to those collected with more frequently used hashtag samples. The results suggest that hashtag samples misrepresent important aspects of Twitter activity and may lead researchers to erroneous conclusions.
Social scientific research analyzing social media is in a period of growth. Studies using big data have appeared in a wide variety of areas ranging from urban sociology (O’Brien, Sampson, & Winship, 2015) to analyses of culture (DiMaggio, Nag, & Blei, 2013). Within this body of literature, social media is perhaps the most widely used source for data collection. Social media, according to boyd and Ellison (2007), is a type of technology emphasizing the connections between users and built to maximize communication among participants in a public or semipublic way. The surge of research analyzing social media is promising and opens up new research questions difficult to answer with conventional samples or statistical techniques.
An important but largely underemphasized aspect of working with databases of social media content is sampling. This may be due to the relative ease of building large databases of user activity, particularly for widely used social media sites such as Facebook or Twitter, coupled with high levels of standardization in sampling procedures due to Application Programming Interfaces (APIs). Sampling social media platforms requires not only platform-specific knowledge of the terminology adopted by users but also which components of the API are more appropriate for a specific research question. In the case of Twitter, analysts typically examine status updates—referred to as tweets—that are brief posts of up to 140 characters linked to a specific user account. Users may embed one or more “hashtags” in their tweets by prefacing a key word with # (e.g., #twitter), post the tweets authored by others (retweets in Twitter terminology), mention other accounts directly (@mentions or @replies), subscribe to the tweets posted by others (followers), or have one’s tweets subscribed to (friends).
Despite the potential of social media platforms as a source of high quality data, two significant gaps remain in scholarly knowledge regarding sampling these technologies. First, attention to sampling is a staple of social scientific research. While most attention focuses on survey research (Alwin, 2007, provides a review), social scientists have studied a wide variety of methodologies. Though examples of research discussing sampling issues for big data can certainly be found (Berinsky, Huber, & Lenz, 2012; Christenson & Glick, 2013), several scholars have noted more research on sampling Twitter in particular is needed (Burgess & Bruns, 2012; Driscoll & Walker, 2014; González-Bailón, Wang, Rivero, Borge-Holthoefer, & Moreno, 2012). Many social scientific studies using Twitter data implement a hashtag sampling design (e.g., Earl, Hurwitz, Mesinas, Tolan, & Arlotti, 2013; Gleason, 2013; Tinati, Halford, Carr, & Pope, 2014), which builds a database by searching the Twitter API for a small number of hashtags (and often just one) present in Twitter status updates. Despite the widespread use of hashtag sampling designs, they have not been systematically evaluated in terms of their coverage and representativeness. Since hashtags provide one of the primary vehicles for data collection, biases in the type of data hashtag samples generate may influence substantive conclusions.
The second gap in the literature is that due to the scale of big data samples, concerns of bias or representativeness may be minimized in a study’s research design or the discussion of findings (González-Bailón et al., 2012). Mythology surrounding big data often incorrectly suggests that bigger data are inherently of better quality (boyd & Crawford, 2012). Sample size alone is not a sufficient metric of sample representativeness. Though the relative ease of collecting millions of cases may provide initial confidence that the data are high quality, it is a misleading way to assess whether a sample covers its target population. Instead, different populations require distinct sampling strategies, and decisions made by researchers about how data are collected directly influence data quality.
This study contributes to the literature on social media–based research by examining sampling procedures for Twitter. I propose strategies for more systematic collection of nonprobability samples using the Twitter API and then empirically validate them in comparison to more widely used hashtag samples. I first propose a set of conceptual distinctions that arise when using Twitter as a data source, differentiating between unbounded populations, semibounded populations, and bounded populations and then outline data collection strategies for the latter three types. Second, I use the proposed data collection strategies to build three defensible representations of the populations of Twitter activity drawn from the Occupy Wall Street (OWS) movement and the Internal Revenue Service (IRS) scandal of 2013. Statistical analysis comparing the populations to hashtag samples reveals stark differences. The results provide evidence that one’s sampling strategy will affect not only the representativeness of the data collected but also the substantive conclusions of analyses.
Nonprobability Sampling and Twitter
Quantitative analyses of social media use a wide variety of sampling techniques for platforms such as Facebook and Twitter (e.g., Gerlitz & Rieder, 2013; Morstatter, Pfeffer, Liu, & Carley, 2013). Big data also bring a unique set of challenges involved in developing a meaningful and representative database of the user activity of interest. boyd and Crawford (2012) discuss how the information used in many big data studies is drawn from commercial third parties, subjecting researchers to the regulations and whims of these organizations. Changes in how a Search API weights and returns data that are not reported in the API documentation, for example, may result in undetected bias in a sample. If data returned from an API are systematically biased, these gaps in the data sets may be unknown and unknowable. Indeed, while populations of Twitter or Facebook data exist in principle, even in cases where researchers purchase data sets from commercial firms, it remains a matter of guesswork for most researchers whether they possess a census or a sample.
A central challenge for big data that is directly related to sampling procedures concerns generalizability. Due to the scale of available data, analysts typically focus on a segment of what can possibly be collected. Researchers working with an API generally do not generally attempt to collect an exhaustive set of Facebook posts or Twitter status updates, instead focusing on a subset of the entire universe of data. Much of this information is collected using nonprobability sampling and often convenience sampling. A careful reading of research using Twitter data indicates that scholars are rarely interested in restricting generalizations to a hashtag sample, which contains all or most tweets using a specific hashtag. Instead hashtags are used as representative proxies of larger conversations on social media. Tremayne (2014), for instance, discusses how Twitter networks influenced OWS’s mobilization patterns generally, even though the sample was based on only two hashtags; other examples can be found in Conover, Davis, et al. (2013), Conover, Ferrara, Menczer, and Flammini (2013), Tan, Ponnam, Gillham, Edwards, and Johnson (2013), and Tinati et al. (2014). The distinction between the sample collected and the goal of broad generalization is quintessentially important, as it can result in a mismatch between whom or what is being sampled, and the scope of any generalizations drawn from subsequent data analysis.
Collecting Twitter samples
Previous research examining sampling has generally produced mixed results about best practices, with scholars calling for more attention to the nuts and bolts of collecting high-quality data (Burgess & Bruns, 2012; González-Bailón et al., 2012). Rieder (2012) proposes a set of distinctions covering different types of Twitter samples, however, does not validate the methodology used to collect data. Gerlitz and Rieder (2013) propose sampling from the 1% of tweets returned by Twitter’s Streaming API, yet acknowledge that this technique may omit entirely topics or accounts that are rare.
To facilitate more systematic inquiry into sampling Twitter, I distinguish between three major types of populations: unbounded populations, semibounded populations, and bounded populations. I subdivide semibounded populations in user-restricted semibounded populations and topic-restricted semibounded populations, resulting in a total of four population types. These distinctions rest on two restrictions that provide boundaries around the sampling frame when building databases of Twitter activity: First, samples may require restrictions on the substantive topic(s) of tweets. Second, researchers may restrict the sample of user activity to only a specific set of user accounts meeting particular criteria. When some combination of these conditions hold, the four types of populations listed above emerge, each with specific considerations for sampling. Table 1 outlines the different types of populations, organized as a typology with distinctions for topical and account restrictions. I now turn to a more detailed discussion of each type of population, with emphasis on semibounded and bounded populations, and propose data collection strategies aimed at maximizing data quality. In certain cases, it is possible to use these strategies to create either populations of activities or pseudopopulations which have close exhaustive coverage, with API constraints.
Typology of Population Types for Twitter Sampling.
Unbounded populations
Unbounded populations are cases where there is no topical restriction on status updates produced by the population of interest along with a simultaneous lack of restrictions on user accounts; that is, unbounded populations include any tweet from any user. While it is possible, in principle, to create a census of activity for unbounded populations, challenges in access to such information are likely prohibitive for nearly all applications that use the Twitter API, though it is possible to purchase such information from commercial firms. In most cases, a random sample of all user activity is the gold standard for building a representative snapshot of unbounded populations. To this end, Twitter provides API tools to randomly sample all public user activity by making subsamples of daily status updates. Unbounded populations are an important area of inquiry for sampling, but beyond the scope of this study for two reasons. First, databases of unbounded populations are best collected via random sampling. I concentrate on nonprobability sampling, which is more widely used in social scientific research. Second, of the research examining sampling issues on Twitter, much of the emphasis has focused on unbounded populations (Driscoll & Walker, 2014; Morstatter et al., 2013).
Semibounded populations
Semibounded populations inform the bulk of substantive research using Twitter. Studies using semibounded populations draw boundaries around the sampling frame, restricting the set of cases of interest based on clearly demarcated criteria. Semibounded populations arise in two ways: First, one can sample on a defined set of users without restrictions on the topical content; second, sampling may focus instead on specific topics without restriction on the set of users. I, respectively, refer to these as user-restricted semibounded populations and topic-restricted semibounded populations. As sampling strategies differ for each type of semibounded population, I discuss each separately.
Studies of user-restricted semibounded populations have focused on a variety of issues. Wu, Hofman, Mason, and Watts (2011) make use of Twitter’s “lists” feature, which lists a specific group of user accounts, to examine how elite users shape the overall tone and direction of online conversations. A growing body of research has also examined Twitter usage patterns for accounts maintained by politicians (Gainous & Wagner, 2013), corporate firms or brands (Kwon & Sung, 2011), or professional athletes (Watkins & Lewis, 2014). Despite the topical diversity, the common thread of these studies is that they focus on the set of tweets produced by specific accounts, and these accounts provide the boundary of the sampling frame.
To maximize data quality for analyses of user-restricted semibounded populations, I propose three steps to facilitate robust data collection. First, produce as comprehensive a list as possible of the accounts of interest. The cultivation of such lists is contingent on the specific research question under investigation, though Twitter features such as lists can be useful, as are databases drawn from sources external to Twitter (e.g., lists of Fortune 500 firms, members of congress, and directories of social movement organizations). Second, collect all tweets posted by each account by making calls to the Twitter API requesting user time lines. Up to 3,200 of the most recent tweets for each account may be collected from the Twitter API, thus multiple rounds of data collection may be required for highly active accounts as new status updates are posted. In cases where tweets beyond the 3,200 most recent posts are required, it may be necessary to supplement the data using a commercial distributor of Twitter data. Third, it is often necessary to oscillate between these steps as new accounts become identified (e.g., from mining the text of the tweets, changes to followers or friends networks, or as new accounts are added to lists). Researchers should therefore actively monitor data collection to ensure that the desired list of accounts remains up to date, and then new accounts are passed to the workflow making API requests for user time lines. Depending on the size and complexity of the accounts under investigation, this data collection strategy may yield either full coverage of the user accounts of interest or a close approximation of all tweets of interest.
Studies examining topic-restricted semibounded populations are also widely used in research. Jones (2014), for instance, examines Twitter discussions on healthcare reform using “#healthcare” as a search parameter. Bastos, Raimundo, and Travitzki (2013) examine gatekeeping patterns in political hashtags, using the case studies of the protests surrounding the 2009 Iranian election, the 2010 popular uprising in Venezuela, and the Egyptian uprising of 2011. Studies have focused on predicting electoral outcomes based on Twitter chatter about political actors (Gayo-Avello, 2013). These studies focus on a small number of topics, but are open to collecting data from any account using a specific hashtag or set of key words, which provide a threshold determining whether the tweet should be included in the analytic sample.
Sampling topic-restricted semibounded populations is the most difficult of the four types of populations identified in Table 1. Unlike the case of unbounded populations, researchers are generally attempting to collect as comprehensive a set of Twitter activity as is possible. This goal is complicated by the fact that the Twitter API is something of a black box in terms of assessing the representativeness of the sample, whether one uses the Streaming API or the Search API, both of which are viable ways to collect topically bounded samples (boyd & Crawford, 2012; Driscoll & Walker, 2014; Gerlitz & Rieder, 2013; Morstatter et al., 2013). Even if a researcher possessed an exhaustive database of all tweets posted during the analytic period of interest (e.g., as one might purchase from a commercial firm such as GNIP), winnowing down the complete Twitter stream to only the content of interest would require the construction of queries to omit tweets irrelevant to the research project which is likely to omit tweets of interest. To maximize the breadth of data collection, multiple queries containing a combination of hashtags, key words, and strings should be passed to the Search API and/or the Streaming API. The search parameters used to collect data need not be mutually exclusive, as redundant tweets may be removed once data collection is complete. For research involving long-term data collection in particular, it is also important to monitor the downloaded data and introduce new search parameters or remove existing ones as the discussion evolves. If one is examining a topic that is trending or otherwise receiving a high level of activity on Twitter, it may be necessary to augment the results from the Search or Streaming APIs with content purchased from a commercial firm. Finally, the decision to use the Streaming API, the Search API, or a combination of the two should be carefully considered. In the case of topics that generate high volumes of discussion, for instance, there is evidence that the Streaming API may provide a more viable sample (Morstatter, Pfeffer, & Liu, 2014; Morstatter et al., 2013). When discussion is relatively limited in its frequency or in applications when the words used to describe a topic are dynamic, the Search API may produce better coverage; however, when the content is generated quickly or using a stable set of key words, the Streaming API may be preferable.
Bounded populations
Bounded populations have both topical restrictions and user account restrictions, making them the most constrained type of population considered here. Since bounded populations are often narrow in nature, they are less widely used in the literature. Several studies have provided important scholarly insights using data on bounded populations: Vis (2013) examines how two prominent reporters covered a series of riots taking place across the United Kingdom, which was linked to a larger sample of Twitter conversations. Bruns, Highfield, and Burgess (2013) analyze interactions between Arabic-speaking, English-speaking, and multilingual Twitter users discussing the 2011 Egyptian uprising and the Libyan civil war.
To maximize data quality for samples of bounded populations, the most effective strategy is to first create a database of the accounts of interest, similar to the proposed data collection strategy for user-restricted semibounded populations. Then, use the Twitter API to collect user time lines for these accounts. As above, multiple rounds of data collection may be necessary as new accounts become relevant or as conversations progress. Next, the results from the user time lines can be pared down so that only tweets matching the topics of interest are retained. Depending on the scale of data collection and the speed at which new data are being produced, this can be done during the data collection process itself by filtering the results, or running each status update through a classification algorithm (e.g., a neural network or support vector machine), many of which are reviewed by Murphy (2012). Alternatively, it is also possible to filter the results after data collection is complete using similar techniques.
Data and Method
I use three data sets to examine Twitter activity for semibounded and bounded populations and to assess coverage patterns for hashtag samples. The core logic of my approach is to first define substantive populations of interest that are consistent with the concepts introduced above and can be approximated using the Twitter API for data collection. Second, I used the strategies above to collect defensible representations of these populations using the proposed data collection strategies. These databases are compared to the less exhaustive, but more widely used, hashtag samples, which are subsets of the larger whole. The data were collected using a customized application making use of the “Python-Twitter” module (Python-Twitter Developers, 2015) in Python.
The first database provides an example of a user-restricted semibounded population. It is based on a sample of 1,359,683 tweets produced by 1,653 accounts maintained by the main organizing groups active in the OWS movement, which represents the population of interest. OWS is a progressive social movement that staged more than one thousand occupations globally starting in the September of 2011 through mid-2012. The analytic period examined here is between July 23, 2011, and May 5, 2012, and captures the main arc of mobilization for the movement, the disbandment of the occupations, and the wave of mobilization for May 1, 2012. I collected all available data on each account once per week starting September 25, 2011, and continuing through May 31, 2012, to overcome the limitations of the Twitter API, which only provides up to 3,200 tweets for each account. The master list of accounts was built by combining directories of Twitter accounts held by OWS groups from three main hubs of activity: http://www.occupy.net/, http://www.occupytogether.org/, and http://www.wealloccupy.com/. I built web crawlers to process the organizational listings on each website and flag new accounts every 24-hr period. Time lines were collected immediately for new accounts, which were then added to the weekly collection queue. To ensure thoroughness once the first wave of data collection was complete, the text for each tweet was processed to locate additional OWS accounts retweeted, @replied, or @mentioned, analogous to snowball sampling. The tweets for the new groups were downloaded, and I repeated this process until reaching complete saturation in terms of new accounts. This process ultimately resulted in a dynamic list of Twitter accounts, with a total of 5,001 accounts that were sampled weekly over the entire data collection period. Once data collection was completed, each account was processed manually to ensure that it was an organizational account explicitly linked to OWS group and active in the movement. This step reduced the final sample to 1,653 accounts, ultimately providing a defensible list of the full population of the OWS movement’s presence on Twitter. All status updates associated with each account were collected with the user time line functionality of the Twitter API. Since data collection began early in OWS’s life cycle, and data on new accounts were collected as soon as they were located, this sampling strategy provided a complete record of all tweets posted by 97.52% (n = 1,612) of the accounts in the analytic sample, often containing coverage beyond the 3,200 tweet limit imposed by the API. Of the remaining accounts for which I did not collect 100% of all tweets, the analytic sample contains a minimum of 75% of all tweets, with a mode of more than 90% coverage. Based on a review of the metadata for each account, this database did not contain 14,233 tweets across all the accounts. Put another way, this database contains nearly 99% of all tweets posted by the 1,653 OWS organizations. Overall, then, this database provides a defensible representation of the population of OWS accounts active on Twitter.
The second data set provides an example of a topic-restricted semibounded population. I use a sample of 814,828 tweets from 155,329 accounts about the IRS scandal of 2013. The sample contains tweets posted between May 20, 2013, and June 5, 2013, which captures the early period of discussion on Twitter. The scandal broke when evidence emerged, suggesting that the IRS was targeting conservative groups filing for 501(c)(4) status, and particularly those groups associated with the Tea Party movement. Several conservative media personalities and Tea Party groups heavily criticized the IRS and Obama administration, and the Tea Party patriots organized 166 demonstrations across the United States on May 21, 2013, to protest IRS oversight.
In this case, the population of interest is all tweets concerning the IRS scandal posted between May 20 and June 5, 2013. I collected data by running multiple queries to the Twitter Search API to database as many tweets on the IRS scandal as possible, given API constraints. In this application, I opted to use the Search API over the Streaming API since the overall flow of content was relatively low, which can result in biased results from the latter API (Morstatter et al., 2013, 2014). As noted above, this decision was based on the ebb and flow of tweets about the IRS scandal. In other applications, the Search API may provide better coverage. Data were collected using a computer program to iterate through a dynamic set of key words making a continuous set of queries to the Search API. The specific parameters included the following: #dcintervention, #irs*, irs, Internal Revenue Service, #tcot, #teaparty, and Tea Party. I then used secondary screening to filter out irrelevant tweets by selecting only those cases that contained the strings “irs,” “internal revenue service,” “monitoring,” or “scandal.” A review using a random sample of 1,000 tweets suggested that this strategy correctly filtered pertinent tweets from the larger database in 98% of cases. As discussed above, developing population databases of topic-restricted semibounded populations is quite complex, even when a research has access to the universe of tweets posted during the analytic period. In this case, a good reason to believe that this database represents either the population of IRS scandal tweets or at least a defensible approximation of it, is the redundancy of tweets returned by the Twitter Search API. That is, my data collection strategy returned a total of 12,195,119 tweets, of which only 6.68% (=814,828) were unique. This indicates that to the degree, my search terms exhausted discussions of the scandal on Twitter—and such a high level of redundancy suggests that this is the case—then this database provides a defensible representation of the population of interest. As well, examining the time stamps of the data suggested that the Twitter conversation about the IRS scandal fell below the Search API limit of 1,000 posts for the entirely of the analytic period.
The final data set represents a bounded population. I use a subset of the full database of OWS groups described above, with the goal of examining all tweets about mobilization posted by the social movement organizations affiliated with the OWS movement. By mobilization, I refer to the thousands of meetings, protests, and other concrete forms of collective action staged by OWS groups. Importantly, this set of tweets excludes nonspecific calls to action or sloganeering; instead, only events and coordinated actions were coded as mobilization events. In other words, this sample restricts both the accounts of interest—only organizational accounts of OWS groups—and the content of the tweets that they produce—only tweets about mobilization. In this case, all tweets concerning mobilization posted between July 30, 2011, and May 5, 2012, were identified using a support vector machine (SVM; see Murphy, 2012). The starting date of July 30 is based on the first tweet about mobilization that occurred in the larger database. The support vector machine accurately identified tweets about mobilization in 92.25% of cases, yielding an analytic sample of 200,601 tweets from 1,547 accounts. The SVM was trained and tested with a random sample of 2,000 tweets that were independently classified by the author and a research assistant. Initial agreement was 97% and cases of disagreement were reexamined and corrected. The recall for the SVM was 85.4% with a precision value of 71.2%, suggesting that the algorithm was highly effective in locating mobilization tweets from the larger database. As noted above, my sampling strategy for OWS groups exhausts the population and provides 100% coverage of a tweet posted for 98% of these accounts. In combination with the high level of predictive accuracy for the SVM, this database provides a rigorous and defensible depiction of the population of mobilization tweets for OWS.
Evaluating Sampling Properties for Twitter Data
My analysis proceeds in two phases, with the goal of first assessing the coverage of hashtag sampling relative to my more systematic approaches that returned populations of tweets. Second, I examine the substantive implications resulting from the different data collection procedures. Since the data collection procedures used here yield what is essentially an exhaustive picture of Twitter activity for the three samples, the statistical estimates are treated as parameters, which can be compared to the statistics estimated from the hashtag samples.
To proceed, I first focus on trends in hashtag usage across the three databases using a Python program that counted hashtags. This allows for direct assessment of whether hashtag samples are biased in their overall coverage of the larger population of user activity. The second component of my analysis examines the substantive implications of sampling procedure. To do so, I use three distinct statistical procedures and compare the results across different data collection techniques. In all comparisons, I used the two most commonly occurring hashtags to construct the latter sample, which is consistent with data collection strategies used in the existing literature (see, e.g., Conover, Ferrara, et al., 2013; Earl et al., 2013; Gleason, 2013; Tan et al., 2013; Tinati et al., 2014). To the degree that the results coincide, the use of hashtag sampling is defensible as a data collection strategy. However, if the results suggest different substantive conclusions, there is evidence that hashtag sampling may provide a biased representation of the larger population of activity on Twitter, despite its widespread use by researchers.
Application 1: User-restricted semibounded populations
To assess hashtag sampling for user-restricted semibounded populations, I examine the intergroup communication patterns for the1,653 accounts identified as core actors in the OWS movement, which I treat as a social network. To measure direct interactions between Twitter accounts, I use either retweets of a status update or @mentions and @replies of accounts in the network. Several scholars have argued that Internet technologies have opened up new avenues for communication among social movements (Bennett & Segerberg, 2013). As a result, the network dynamics of Twitter discussions can show how social movement actors and organizations collaborate or exchange information. I wrote a Python program to extract the discussion patterns from the larger database by identifying all retweets, @mentions, and @replies. This information was used to construct a directed graph counting the number of times each account communicated with the other accounts over the analytic period. A second network was also built based on a hashtag sample containing all tweets with #occupy or #ows—the two most commonly used hashtags in this sample (see Table 2). I then examined and compared the basic structural characteristics of both networks to see whether the results were consistent across the two samples.
Prevalence of the 10 Most Commonly Used Hashtags.
Note. Percentages are based on hashtag usage across tweets and repeated hashtags within a single tweet are counted once. All hashtags are treated as case-insensitive for calculations. IRS = Internal Revenue Service.
Application 2: Topic-restricted semibounded populations
The second application focuses on topic-restricted semibounded populations, using the data set on the IRS scandal as a case study. I examine patterns of leadership and influence on Twitter by looking at the most retweeted accounts in the sample, comparing the results from the full sample to a hashtag sample containing all tweets with #irs and #tcot. As in the first application, these are the two most commonly used hashtags in the data (see Table 2). Over, 40% of the full database consists of retweets, with the vast majority of retweets emanating from messages posted by a small number of conservative politicians, news organizations, and media personalities. These individuals and organizations emerged as leaders in the backlash against the IRS and played a critical role in shaping the tone and scale of the conversation, making it important to identify who is being retweeted, and how frequently. I developed a Python program to process the data set and identify retweets and the accounts being retweeted, which is the central variable used in the analysis below.
Application 3: Bounded populations
The final application focuses on bounded populations using group-level mobilization processes for OWS. The potential for mobilization on social media has been treated as perhaps the most revolutionary change brought about by internet technologies. With a small number of keystrokes, it is possible, at least in principle, to reach millions of individuals and plan events in real time. It follows that studies examining mobilization patterns online require a comprehensive picture of how social movements use platforms such as Twitter or Facebook to organize events. Systematic gaps arising from sampling design may underestimate or overestimate the scale of mobilization organized online. To assess trends in online mobilization, I compared the full database of OWS mobilization tweets to the hashtag sample returned by using the search parameters #ows and #occupy, the two most common hashtags. I counted the number of mobilization tweets occurring in the full and hashtag samples for each account, which provides a concrete way to assess the degree to which the latter sample provided good coverage of the larger population of Twitter activity.
Results
Usage Patterns of Hashtags
The coverage patterns for the hashtag samples relative to the full samples point to a sharp contrast. For the user-restricted semibounded OWS database, approximately 20% of tweets are retained in the hashtag sample. Comparable figures are, respectively, 33% for the IRS data and nearly 20% for the bounded OWS data. While the relative uniformity of the disparities in coverage across the three types of samples is encouraging, it is important to emphasize that in all cases, a significant percentage of the pertinent user activity is omitted from the sample. Unless a compelling argument can be made that the 20% of tweets returned via hashtag sampling do not differ in content from the larger population, there is a strong initial case for bias.
Figure 1 extends this discussion and summarizes the distribution of hashtags across each data source. The modal number of hashtags used in tweets is zero. This is true in approximately 45% for the user-restricted semibounded population, 60% for the topic-restricted semibounded population, and 40% of cases for the bounded population. Consequently, even samples collected by searching for the most widely used hashtag could potentially overlook approximately half of all user activity. Examining the larger distribution of hashtag usage across the three databases, there are strong patterns of commonality. A more in-depth analysis of the results in Figure 1 indicates that between 90% and 91% of tweets have three or fewer hashtags across the samples. To the degree that these patterns generalize beyond the three data sets analyzed here, there is evidence of a significant coverage gap arising from samples collected using a small number of hashtags. The data collection strategies I proposed do not have this limitation, as data collection is not subject to usage patterns of hashtags alone.

Hashtag usage patterns across three types of Twitter populations. (a) User-restricted semi-bounded population: Occupy Wall Street (n = 1,359,683), (b) Topic restricted semi-bounded population: IRS Scandal (n = 814,828) and (c) Bounded population: Occupy Wall Street mobilization sample (n = 200,601).
The results in Table 2 demonstrate that no more than 30% of posts contain the modal hashtag. For the semibounded OWS database, #ows appears in just 18% of tweets, while for the IRS data, #irsis present 29% of the time. In the bounded OWS database, #occupy is in 18% of tweets. Of particular interest for the IRS data is that the Tea Party patriots—a national umbrella group of Tea Party organizations—organized a nationwide series of protests on May 21, which was publicized using the suggested hashtag #dc intervention. This hashtag appeared in just 0.06% of tweets. Sampling on this hashtag alone would clearly lead to only a small segment of the larger discussion of the IRS scandal. Last, even though the semibounded and bounded OWS data sets are drawn from the same set of accounts, the most common hashtag differs across the two samples. This difference could be quite consequential for data collection if only one of the two hashtags was used for data collection.
The Empirical Consequences of Sampling Strategy
Having examined a baseline comparison of hashtag usage for the large data sets of Twitter activity, I now turn to three empirical examples with the goal of comparing the statistical results returned by hashtag samples to their more comprehensive population databases. This allows me to assess whether hashtag samples provide more representative and bias-free coverage of certain types of populations. In each application, the full database of tweets is compared to its corresponding hashtag sample, which is based on the two most widely used hashtags from Table 2.
Application 1: The network dynamics of Occupy Wall Street on Twitter
I begin with a network analysis of communication patterns between accounts linked to the OWS movement to study user-restricted semibounded populations. The results are presented in Table 3.There are significant differences in the structural dynamics of the networks estimated using the full population of tweets and a hashtag sample. The latter consistently misrepresents central aspects of OWS’s discussion network. For example, in the full database, there are only 11 accounts isolated from the larger network, yet for the hashtag sample, there are 167 isolates, a more than 1,400% increase. The connectivity of the network is underestimated not just in density, which is expected given the necessarily smaller sample size of the hashtag sample, but also in identifying connections between nodes in the network. The population database contains 55,820 unique ties between accounts while the hashtag sample suggests that there are only 22,055. This is further evinced by comparing the average size of the K-cores, which measure the degree to which subcomponents of the graph are connected, as the hashtag sample returns a value of approximately 16, while the full sample has a comparable estimate of 36. This suggests considerably lower levels of connectivity in the network built using the hashtag sample. Turning to the measures of centralization and variability in in-degree and out-degree ties, we again see notable differences across the two networks. Degree centralization for the hashtag sample is 0.31, while for the population network, the value is 0.38. In the case of the standard deviations for both in-degree and out-degree ties, the hashtag sample returns a consistently smaller value than the population. There is a downward bias in the estimates for whether retweets and mentions are reciprocated in the hashtag sample, with values of 0.29 and 0.20 for the full and hashtag data, respectively.
Network Characteristics of Replies on Mentions on Twitter for Occupy Wall Street.
Note. Estimates are based on 1,359,683 tweets in the full sample and 271,979 tweets in the hashtag sample. The hashtag sample uses #ows and #occupy as search terms. SD = standard deviation.
Application 2: Influence on Twitter and the IRS scandal
I next turn to an analysis of a topic-restricted semibounded population by examining the database on the IRS scandal. Table 4 outlines the 10 most retweeted accounts for the full sample and hashtag subsample. This is used to assess who was leading the Twitter discussion and the degree to which conclusions about the most influential accounts are sensitive to sampling strategies. Results from the two databases are quite distinct in several ways. In the population of tweets, 23,851 accounts were retweeted a total of 457,308 times. For the hashtag sample, 9,149 accounts were retweeted 167,424 times. That is, for each retweet in the hashtag sample, there are about 2.7 retweets in the full sample, pointing to a significant coverage gap in the data.
Ten Most Retweeted Accounts Discussing the IRS Scandal.
Note. Accounts maintained by individuals who are not politicians are anonymized for privacy. The hashtag sample uses #irs and #tcot. IRS = Internal Revenue Service.
The undercounting is to be expected since the hashtag sample is lower in size. However, the results in Table 4 indicate that even after standardizing for the number cases, there are notable shifts in the influential actors across the two databases. Five of the most retweeted accounts are not present in the top 10 for the hashtag sample. There are also differences in the type of accounts represented. In the population, three of the top 10 accounts are media outlets (@breitbartnews, @sharethis, and @druge_report), yet only @foxnews is present in the hashtag sample. Further, politicians are overrepresented in the hashtag sample, appearing three times (@sentedcruz, @darrellissa, and @marcorubio), while only one politician, Ted Cruz, was heavily retweeted based in the population. Both databases have six media personalities that are highly retweeted, though only 66% present in the population are listed in the hashtag sample.
Table 4 points to the sensitivity of hashtag sampling in another way: The most retweeted account in the population, @breitbartnews, does not appear in the top ten for the hashtag sample. This account is instead ranked 22nd with only 786 retweets. If a tweet that becomes heavily retweeted happens to omit the hashtag(s) used for data collection, the discussion will be missed entirely. In this case, the coverage gap was sufficiently large that the influence of the most highly retweeted account in the population would be misstated.
Overall the hashtag sample provides a poor representation of the most influential actors during the IRS scandal. This is not just in terms of the frequency of retweets, but also in terms of who led the conversation. The hashtag sample overstates the role of politicians and underestimates the impact of news outlets, two of which are leading aggregators that lean conservative. If the hashtag data were used in a more developed analysis of the IRS scandal, biases in the data may significantly alter the substantive conclusions of the research.
Application 3: OWS and mobilizing on Twitter
The final application is for a bounded population and uses the set of tweets about mobilization produced by 1,547 accounts affiliated with the OWS movement. The central goal here is to assess the percentage of tweets about mobilization is contained in the hashtag sample at the account level. As in the case for both types of semibounded populations, the hashtag sample contains a biased representation of the population. On average, accounts in the population posted 130 tweets about mobilization, while in the hashtag sample, the comparable mean was just 26. The hashtag sample also underreports the number of accounts that engaged in any mobilization on Twitter. Here the hashtag sample suggests that an additional 325 accounts did not post any status updates about mobilization, yet in the full sample, these accounts posted between 1 and 327 relevant tweets.
Figure 2 provides a histogram detailing the coverage of the hashtag sample at the account level for the 1,547 groups. The distribution has a strong positive skew. These results point to a significant problem in coverage for the hashtag sample. The hashtag sample contains fewer than 20% of the mobilization tweets for more than 60% of the accounts and has less than 50% coverage for more than 90% of accounts. These differences remain when the frequency of tweets is considered. For accounts with at least 100 status updates about mobilization, the hashtag sample contains 18% coverage. For accounts with at least 500 tweets, coverage is just 20%, and finally, for accounts with 1,000 or more posts, just 22% were captured in the hashtag sample. Finally, Figure 2 indicates the hashtag sample does have a high level of coverage for a small subset of accounts. However, these accounts typically posted few status updates about mobilization: the median of number of tweets for accounts with at least 50% coverage is 11, while for 80% coverage, the median is 6.5. Thus, the appearance of good coverage in the hashtag sample is largely explained by the infrequency of mobilization tweets.

Histogram of group-level percentages of coverage for tweets about mobilization in the Occupy Wall Street movement. Note. Estimates are based on 1,547 groups.
To summarize, the hashtag sample provides a problematic representation of how OWS groups used Twitter as a platform for mobilization. The hashtag sample is inconsistent in its coverage and only provides a reasonable representation of the scope of Twitter activity for the least active accounts. Since understanding the degree to which social movement organizations use social media technologies to mobilize has emerged as a central problem in social movements research (Bennett & Segerberg, 2013), systematic biases in samples may unnecessarily slow scientific advancement and lead researchers to incorrect conclusions.
Discussion and Conclusion
This research had two major goals: First, I developed a typology outlining four general types of populations that emerge for studies using Twitter data and proposed sampling strategies linked to each; second, I validated this typology using three large data sets of Twitter activity intended to provide defensible representations of the populations of interest, up to the limits of the Twitter API. The goal was to assess and compare the data collection strategies I propose to hashtag samples, which are more widely used. The overall results suggest that hashtag samples are insufficient for building accurate representations of Twitter activity, excluding the relatively rare case in which the aim is to generalize only to the tweets using the hashtag(s) of interest. My typology addresses these limitations and proposed a set of concepts and sampling strategies that are grounded in observable research practices in the social sciences that use Twitter data.
The sampling strategies proposed in this study can be implemented with existing resources provided by the Twitter API, and though they may not exhaust every imaginable scenario where one would sample Twitter, they are applicable for a wide variety of topics and research projects that capture what researchers are concretely working on. In the cases examined here, the more systematic and comprehensive database can generate databases that approximate the full population of interest well. The OWS samples, for instance, omitted only 14,233 tweets for the population of 1,653 account of interest. It is possible, of course, that these tweets may slightly alter the findings here, but since the full database contained 1.36 million accounts, dramatic changes are unlikely. In the case of the IRS scandal data, it is more difficult to assess coverage due to the black box of the Twitter API (boyd & Crawford, 2012). Even purchasing the universe of tweets posted during the analytic period of the IRS data would pose difficulties, as the scale of the information requires algorithmic extraction of the pertinent tweets, which would almost certainly omit a subset of the pertinent content. It is, therefore, important to emphasize that the results for the topic-restricted semibounded population are tentative and contingent on future research.
I close by discussing two implications raised by this research, and then turn to how the analyses here may inform future research. First, it is important to discuss the empirical basis for how hashtags structure conversations on Twitter. Hashtags are useful in that they can provide organizing devices for exchanges on Twitter. However, in each of the case studies examined above, the modal tweet contained no hashtags, and even the most widely used hashtags were present in between 17% and 28% of status updates. This is indicative that even if hashtags have the potential to provide a unified way to engage with other users, they are not used in such a manner in practice. Conversations do not revolve around hashtags, based on the evidence here, and instead are dynamic in content as new hashtags are added or dropped as conversations evolve. Further, Twitter activity is not wedded to a small number of topically specific hashtags. In the case of the IRS sample, the second most widely used hashtag was #tcot. This is well beyond the comparatively narrow scope of #irs, yet #tcot contained a significant component of the discussion about the IRS scandal. Sampling strategies should reflect these dynamics if the goal is to collect a representative snapshot of user activity.
A second implication is that large-scale data collection alone is insufficient to eliminate biases introduced by research design and sampling. boyd and Crawford’s (2012) cautionary arguments about the mythos of big data are quite important to consider in this regard. In each of the three applications, the larger empirical patterns were misrepresented in the hashtag sample, and sometimes, the differences were quite dramatic. It follows that the central point of concern is how hashtag samples misrepresent underlying statistical distributions and facilitate erroneous conclusions due to sampling bias. Hashtags are certainly important components for rigorously sampling Twitter but do not exhaust the steps needed to build high-quality databases of activity. Hashtag searches with wildcard characters may increase coverage if the hashtags are relatively uniform in their composition, such as work done by Conover, Davis, et al. (2013), yet this approach works best in the rare case that the hashtags are quite similar at the onset.
To conclude, there are several points raised in this study that may inform future research. First and foremost, less systematic approaches to sampling are insufficient for capturing the dynamics of Twitter activity. Sampling considerations should become central components of a project’s research design, and the strengths and limitations of different sampling strategies should be explicitly included in the discussion of results. Second, the three case studies I examine focus primarily on social movements, it is likely that the conclusions here generalize to other topics. This remains to be seen, and future research may address whether the degree of bias uncovered here increase or decrease in other substantive applications. Finally, similar to all case studies, generalization beyond the subject population requires caution; however, the conceptual distinctions for different populations on Twitter, and the corresponding sampling strategies I proposed for each, are more general in scope. Scholars have noted that one of the most significant limitations to big data sampling is a lack of consistent terminology (Bruns & Burgess, 2012; Driscoll & Walker, 2014; González-Bailón et al., 2012). This is a consequential omission, as the absence of a crisp, widely applicable set of concepts encourages errors in generalization. Future research can make use of these distinctions to help refine and extend the sampling strategies proposed here, particularly for topic-restricted semibounded populations, and more generally build effective sampling strategies for research using Twitter.
Footnotes
Authors’ Note
This study was presented at the 2014 Meetings of the American Sociological Association, New York, NY. I thank David Garson, the anonymous reviews, and Katherine Johnson for feedback.
Declaration of Conflicting Interests
The author declared no potential conflict of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was funded by the National Science Foundation (SES-1321802).
