Abstract
Twitter is a communication platform that can be used to conduct health science research, but a full understanding of its use remains unclear. The purpose of this narrative literature review was to examine how Twitter is currently being used to conduct research in the health sciences and to consider how it might be used in the future. A time-limited search of the health-related research was conducted, which resulted in 31 peer-reviewed articles for review. Information relating to how Twitter is being used to conduct research was extracted and categorized, and an explanatory narrative was developed. To date, Twitter is largely being used to conduct large-scale studies, but this research is complicated by challenges relating to collecting and analyzing big data. Conversely, the use of Twitter to conduct small-scale investigations appears to be relatively unexplored.
Twitter is an information network that was established in 2006, and it has been quickly adopted by Internet enthusiasts (Twitter, 2014a, 2014b). Currently, it is estimated that 19% of online adults use Twitter (Brenner, 2014). The network can be accessed free-of-charge from a desktop or a handheld device, and it is supported by many mobile phone and tablet applications (i.e., apps). Unique to Twitter is the ability to publicly send and receive brief messages in real time. Given the potential to instantaneously reach untold numbers of people, Twitter is viewed as a relatively unexplored resource for conducting health science research (Yoon, Elhadad, & Bakken, 2013).
With this in mind, the purpose of this systematic narrative review of the literature was to examine how Twitter is currently being used to conduct research in the health sciences and to consider how it might be used in the future. Background information relating to Twitter is discussed next, followed by an explanation of the methods that were used to conduct this review, the review findings, and a discussion.
Background
Twitter users and use. Twitter users include individuals and organizations. To access Twitter, a person or organization must establish an account. Accounts can be designated as public or private, but unlike platforms such as Facebook, the default setting is public (Lyles, López, Pasick, & Sarkar, 2013). Each account holder has a profile page, which generally features a customized background, picture, web address, and brief biography (Lulic & Kovic, 2013). Many account holders (~40%) read messages but do not post, and those who post can have multiple accounts (Twitter, 2011).
After establishing an account, Twitter activity is initiated and expanded by posting messages (i.e., tweets), reposting noteworthy messages (i.e., retweeting), and attracting other individuals (i.e., followers) to the account. Account holders establish customized message feeds or streams by applying selection criteria, which are commonly based on keyword searches. Once an account holder’s interests are known, Twitter will algorithmically suggest accounts to follow (Lyles et al., 2013; Twitter, 2014c).
Tweets are limited to no more 140 characters, and they are referred to as microblogs. Within the character limitation, tweets can include links to webpages, and they can also include pictures. Public tweets are accessible to any Twitter user, but they can also be sent to targeted followers or individuals who are not followers. With a user’s permission, each tweet can be tagged so that its geographic origin is known (Twitter, 2014c). This feature, with the help of other technological add-ons, has the capacity to produce sizable samples of valuable geo-located Twitter data (Weidemann & Swift, 2013).
The extent that a message is distributed depends on the number of followers who are linked to an account and the number of times that a message is retweeted (Twitter, 2014c). All Tweets are posted immediately, thus, enabling real-time feedback (King et al., 2013; Twitter, 2014c). Tweets appear in list format on an account holder’s home page, with the most recent messages on top. This chronological display constitutes what is known as a Twitter timeline (Twitter, 2014c).
Twitter reports 255 million active users per month and 500 million Tweets per day. The majority of account holders are located outside of the United States (77%), and 78% of them access the communication network from mobile devices (Twitter, 2014a). From a research perspective, text-based tweets are considered data as well as the metadata that accompany them. Metadata include each account holder’s language, their geo-location, the number and names of the people they follow, as well as the number and names of their followers (Mayer-Schonberger & Cukier, 2013).
Of some concern for investigators is the fact that a small percentage of private Twitter accounts and/or individual private tweets are not accessible. Although dismissed as insignificant by some, private messages are likely to reflect perspectives that are not shared in public posts; thus, some viewpoints might be lost to analysis (boyd & Crawford, 2012; Hajar, Clauson, & Jacobs, 2014; Lyles et al., 2013). This is a particular concern for researchers who are interested in the perspectives of marginalized individuals who do not feel comfortable posting their thoughts and opinions in public fora (Umihara & Nishikitani, 2013).
Account holders have unrestricted access to their own account metadata and tweets, which they can download in Excel format, but special permission is required to gain full access to the Twitter platform. Currently, individuals can download a limited amount of public data without special permission and/or a fee (Boulos & Anderson, 2012; boyd & Crawford, 2012; Mayer-Schonberger & Cukier, 2013). Notably, individuals who download data from the Twitter platform are unlikely to have full insight into what portion of the database they are accessing (e.g., random, cross section, most recent, etc.), and thus, the context and generalizability of their findings remain unclear (boyd & Crawford, 2012). Currently, the rules and mechanisms for accessing the Twitter database are in flux, and they are highly technical. Investigators who are interested in downloading data are urged to consult the Twitter website (https://dev.twitter.com/streaming/overview) for the most up-to-date information.
The above challenges aside, in 2010, Twitter.com donated its entire archive of public Tweets (past and future) to the Library of Congress (LOC) for purposes of preservation and non-commercial research. To date, the Library has been preparing the archive for use and has received hundreds of inquiries about access. So far, the archive is unavailable (LOC, 2013), and access guidelines remain unclear.
Several research institutions have also been granted full access to the Twitter database. They include the Massachusetts Institute of Technology, Harvard Medical School, and Boston Children’s Hospital. The latter two are currently using Twitter messages to study the spread of foodborne illnesses (Bray, 2014).
Big data
For individuals who use social media solely to stay apprised of current events and to communicate with family and friends, it might be difficult to envision the amount and type of data that are and will be available from the Twitter platform. Using the current vernacular, we are looking at big data. Simply put, this means data in quantities so large that investigators must develop new ways to manage and analyze it (boyd & Crawford, 2012; Mayer-Schonberger & Cukier, 2013).
Given the volume of data that social networks such as Twitter offer, it is argued that sampling error is diminished along with the need for random sampling. Also, as sample size increases, emergent patterns are easier to identify, and researchers are more interested in correlations and probabilities than causal analyses. As such, researchers are trading insight into why something happens for insight into what is happening. In general, attention is being focused on macro-trends and patterns versus micro-perspectives (Mayer-Schonberger & Cukier, 2013; Myslín, Zhu, Chapman, & Conway, 2013).
Using a big data lens, investigators do not clearly articulate research questions at the outset of investigations. Rather, they mine the data and inductively infer research questions as patterns emerge. For instance, as trends become apparent, investigators might formulate questions regarding a specific subgroup of individuals and, in turn, pursue more focused analyses (Mayer-Schonberger & Cukier, 2013).
Despite considerable enthusiasm for big data, large datasets come with challenges. One challenge is that data extracted from Twitter are likely to contain considerable error. This is due to a number of factors, including the fact that Twitter is riddled with computer-generated spam as well as messages that are filled with topically irrelevant information, unorthodox abbreviations, and misspelled words (Lyles et al., 2013; Montejo-Ráez, Martínez-Cámara, Martín-Valdivia, & Ureña-López, 2014). To make matters worse, the Twitter archival and search functions are relatively unsophisticated (boyd & Crawford, 2012), which makes focused data retrieval difficult. To overcome these challenges, self-correcting computer algorithms are being developed to capture, clean, analyze, and visualize data and resultant findings (Mayer-Schonberger & Cukier, 2013).
Another challenge posed by data that are extracted from Twitter is that they are not necessarily representative of the general population. For example, journalists have identified that about half of Twitter users are young, well educated, mobile adults whose views do not necessarily match the general public’s views (A. Mitchell & Guskin, 2013). Moreover, it is feared that powerful entities, such as government agencies and businesses, will inappropriately link findings from Twitter-based research to under-represented groups (Mayer-Schonberger & Cukier, 2013). This is a problem that could be exacerbated by spurious correlations that commonly emerge when dealing with massive amounts of data (boyd & Crawford, 2012).
Paradoxically, while individuals fear over generalization of findings, they also worry about the loss of anonymity and privacy because their every tweet is archived and potentially analyzed (boyd & Crawford, 2012; Mayer-Schonberger & Cukier, 2013). On one hand, Twitter account holders understand that tweets are public. On the other hand, they perceive that ethical parameters should apply to their use in research. For this reason, there have been calls for more accountability in terms of the use of public tweets for research purposes (boyd & Crawford, 2012).
Little data
Although the topic of big data is currently getting a lot of attention, experts agree that researchers should not lose sight of little data (boyd & Crawford, 2012; Mayer-Schonberger & Cukier, 2013). This is especially true because small datasets can easily be extracted from researcher-developed Twitter accounts, and highly contextualized insights can be used by healthcare providers to promote personalized health and well-being. For example, based on Twitter communiques, healthcare providers could potentially provide feedback to individuals regarding their emotional patterns, levels of physical activity, and number of visits to fast-food restaurants (Bonchek, 2013).
Big data and little data aside, data collection and analysis of all kinds can be challenging and expensive. Thus, the hope is that communication platforms such as Twitter will make these tasks less costly, time intensive, and laborious (boyd & Crawford, 2012). To learn more about how Twitter is currently being used to conduct health-related research and to infer how it might be used in the future, a systematic narrative review of the health sciences literature was conducted.
Method
First, PubMed was searched to identify article titles that include the keywords Twitter and tweet. This search resulted in 56 articles that were published between May 2012 and May 2014. To ensure that reports of nursing research were captured, the same keywords were used to search the Cumulative Index to Nursing and Allied Health Literature (CINAHL) database during the identical time period. This second search yielded 12 unique citations for a total of 68 articles.
The 2-year search timeframe (i.e., May 2012-May 2014) seemed reasonable given that Twitter was not in use prior to 2006, and the characteristic organizing hashtag format (e.g., #ruralhealth) was not introduced until 2007. Moreover, because research relating to communication platforms is changing quickly, limiting the search to the two most recent years minimized the threat of accessing obsolete information.
All 68 references were carefully reviewed, and articles in which Twitter was not the focus of health-related research were excluded. In addition, articles were excluded if they were not written in English or if the information relating to Twitter could not be differentiated from that pertaining to other communication platforms such as Facebook. In the end, 31 research articles were identified and secured for review.
Each article was carefully read, and important information was highlighted. Thereafter, notes were taken and categorized based on the research topics, methods (i.e., data collection, samples, and analysis), and findings. Memos were subsequently written and refined until comprehensive and cohesive descriptions relating to each category were developed. These refined descriptions were then compiled to create the following review and discussion of the literature.
Results
It is clear that researchers from around the globe (e.g., United States, Canada, United Kingdom, Japan, and Korea) are curious about how Twitter can be used to investigate health-related issues. In addition, it appears that government and private funding agencies are interested in the same thing, because 15 of the 31 studies that were reviewed were supported by such grants. Research topics, data collection strategies, data analysis methods, and findings relating to Twitter-focused investigations are presented next.
Research Topics
Research topics that were examined generally fell into three categories: health issues and problems, health promotion, and professional communication. Health issues and problems are discussed first.
Health issues and problems
Researchers have used Twitter to examine users’ organic commentaries regarding health problems such as cancer (Himelboim & Han, 2014; Sugawara et al., 2012), dementia (Robillard, Johnson, Hennessey, Beattie, & Illes, 2013), acne (Shive, Bhatt, Cantino, Kvedar, & Jethwani, 2013), cardiac arrest (Bosley et al., 2013), and tobacco use (Myslín et al., 2013). They have also studied in vivo communiques relating to prescription drug abuse (Hanson, Cannon, Burton, & Giraud-Carrier, 2013); and they have specifically looked at geo-tagged tweets pertaining to Adderall misuse on college campuses (Hanson, Burton, et al., 2013). Geo-tagged tweets have also been used to examine pockets of psychological well-being (i.e., happiness) within the United States (L. Mitchell, Frank, Harris, Dodds, & Danforth, 2013).
Health promotion
Researchers have also examined commentaries pertaining to health promotion topics such as smoking cessation (Prochaska, Pechmann, Kim, & Leonhardt, 2012), dietary habits (Hingle et al., 2013), fitness (Vickey, Ginis, & Dabrowski, 2013), vaccinations (Love, Himelboim, Holton, & Stewart, 2013), and cancer screening (Lyles et al., 2013). Moreover, they have investigated how Twitter is currently being used to communicate health promotion messages (Donelle & Booth, 2012) relating to diabetes (Harris, Mueller, Snider, & Haire-Joshu, 2013) and health literacy (Park, Rodgers, & Stemmle, 2013). In terms of health policy, researchers have examined how Twitter communiques might influence political processes (King et al., 2013; Yamaguchi et al., 2013).
Investigators have also looked at the capacity of Twitter to disseminate information about communicable disease transmission and crises situations in real time (Kim, Seok, Oh, Lee, & Kim, 2013; Nagel et al., 2013; Umihara & Nishikitani, 2013). In addition, they have examined factors that might influence the impact of these types of messages, including their source (e.g., expert, non-expert, corporation, computer [bot-controlled]), distance from their source, and the number of source followers (Lee & Sundar, 2013; Tavares & Faisal, 2013).
Professional communication
Researchers have been interested in knowing how healthcare professionals such as pharmacists, emergency physicians, and kidney specialists use Twitter to communicate (T. Desai et al., 2012; Hajar et al., 2014; Lulic & Kovic, 2013; Nomura, Genes, Bollinger, Bollinger, & Reed, 2012). They have also examined how Twitter can be used to provide evaluative feedback to students in clinical settings (B. Desai, 2014).
Finally, individuals who are involved with health sciences publishing have been interested in knowing how Twitter can and should be used to promote professional journals (Boulos & Anderson, 2012). In addition, they have been interested in knowing whether citation impact scores are consistent with Twitter altmetrics (Thelwall, Haustein, Larivie’re, & Sugimoto, 2013).
Data Collection Strategies
To conduct the studies listed above, data were either extracted from the Twitter platform or they were downloaded from researcher-created accounts. In the case of researcher-created accounts, investigators recruited followers to accounts to observe how they organically used Twitter in circumscribed situations. In addition, investigators recruited followers to accounts and then analyzed how Twitter could be used to accomplish specific objectives (e.g., T. Desai et al., 2012; Hingle et al., 2013; Lee & Sundar 2013; Nomura et al., 2012). Among the reports that were examined, collecting data from researcher-created accounts was more the exception than the rule (e.g., B. Desai, 2014; T. Desai et al., 2012; Hingle et al., 2013; Lee & Sundar, 2013; Nomura et al., 2012).
Far more common than gathering data from researcher-created accounts were situations in which investigators set their sights on accessing large samples based on keyword searches of public tweets. Searching was carried out using the Twitter search function or by using custom-designed software. Due to the volume of data available, these searches were generally date restricted (e.g., King et al., 2013), and sometimes researchers used other selective or random sampling techniques (Donelle & Booth, 2012; Hajar et al., 2014; Hanson, Cannon, et al., 2013; Himelboim & Han, 2014; Park et al., 2013; Robillard et al., 2013). In addition, data were filtered according to study criteria such as names, locations, web addresses, account creation dates, number of followers, number of tweets and retweets, and so forth (e.g., Hajar et al., 2014; Love et al., 2013; Lulic & Kovic, 2013).
Based on the articles reviewed, it is apparent that development of customized software for purposes of collecting data from the Twitter platform is a popular enterprise, with no less than seven programs (e.g., Tweet Archivist, Donelle & Booth, 2012; followerwonk, Hajar et al., 2014; ViBE, Hingle et al., 2013; Twitter4J, Kim et al., 2013; Twiangulate, Lulic & Kovic, 2013; Bettween, Sugawara et al., 2012; TwapperKeeper, Vickey et al., 2013) mentioned. In fact, an important objective of some investigations was software development (e.g., Creepy Crawler, Tavares & Faisal, 2013).
Data Analysis Methods
In general, data analysis methods were in keeping with the exploratory nature of the research questions that were asked. Descriptive statistics were calculated in regard to the number of account followers, tweets, retweets, and so forth. When samples were relatively small, manual data categorization methods were used to analyze text-based posts (e.g., King et al., 2013; Lyles et al., 2013; Park et al., 2013). In the case of larger datasets, computer algorithms were often used. For example, in several instances, computer software was used to classify tweets according to sentiment (e.g., happy, positive, neutral, negative; T. Desai et al., 2012; King et al., 2013; L. Mitchell et al., 2013; Myslín et al., 2013).
To search for patterns within large datasets, computer-learned algorithms were used (e.g., Support Vector Machine algorithms and naive Bayes classifiers, Myslín et al., 2013), and customized visualization software was used to display the findings (e.g., mentionmapp, Sugawara et al., 2012; NodeXL, Lulic & Kovic, 2013; NVivo; Harris et al., 2013; GMap, Hingle et al., 2013). In the end, the goal of these types of big data investigations was to develop generalizable models (Mayer-Schonberger & Cukier, 2013).
Findings From Twitter-Focused Investigations
Health-related messages
Based on the studies that were reviewed, it appears that, outside of spam, three types of health-related messages predominate on the Twitter network. The first are akin to commentaries, arguments, or opinions, such as the kind that might be distributed by special interest groups (e.g., be an organ donor). The second type tends to be highly personal and represents moment-to-moment sentiments and emotions (e.g., feeling good today). The third type is broadly categorized as informational (e.g., vaccines fight disease), and it might be promulgated by businesses (e.g., drugstores) or healthcare agencies (e.g., Centers for Disease Control and Prevention).
In keeping with others’ conclusions, Twitter seems well suited to disseminating informational messages because the platform functions more like a public broadcasting system rather than a tool for social communication (Brenner, 2014; Neiger, Thackeray, Burton, Giraud-Carrier, & Fagen, 2013). Given this insight, researchers who are seeking a platform to study social interaction might want to consider other sites such as Facebook, especially because it is oriented more toward privacy.
Sources of Twitter messages
Among the studies reviewed, Twitter messages of all kinds tended to originate from two general sources, the general public and healthcare organizations and professionals. Messages originating from the general public are discussed first, followed by a discussion of those associated with health-related organizations and professionals.
It appears that laypersons can use Twitter to learn and share timely information about health-related issues, problems, and experiences (Bosley et al., 2013; Himelboim & Han, 2014; Lyles et al., 2013). It also appears that laypersons are able to help each other manage challenging health problems such as cancer (Sugawara et al., 2012). Moreover, some individuals are willing to post personal and intimate information regarding health-related experiences (Lyles et al., 2013).
In terms of health-related habits, it is possible to capture the public’s thoughts and behaviors regarding food and eating (Hingle et al., 2013), exercise (Vickey et al., 2013), and smoking (Myslín et al., 2013). In addition, the general public and advertisers are using Twitter to promote smoking cessation; however, neither the public nor advertisers are consistently promulgating evidence-based strategies. Moreover, sustained use of these types of Twitter feeds by the general public appears to be limited (Prochaska et al., 2012).
Based on tweets from the general public, it is possible to identify and characterize prescription drug abuse activities, including where such activities are occurring (Hanson, Burton, et al., 2013; Hanson, Cannon, et al., 2013). Happiness sentiments can also be successfully geo-located (L. Mitchell et al., 2013), and it is possible to quickly track the spread of communicable diseases such as the flu (Kim et al., 2013; Nagel et al., 2013). Twitter users can also stay informed about natural disasters in real time, but unfortunately, not all of the information that is shared during crises is credible. Thus, Twitter-based communiques do not always help to resolve problems and reduce distress (Umihara & Nishikitani, 2013).
Fortunately, laypersons seem to make critical judgments about the credibility of tweets and retweets based on their original source. For example, they view messages from healthcare organizations as more credible than those from laypersons (Lee & Sundar, 2013; Love et al., 2013). In addition, when celebrities tweet about health-related topics, their notoriety versus the content of their tweets appears to have the greatest impact (Himelboim & Han, 2014).
Healthcare organizations appear to be able to educate the public using Twitter (Bosley et al., 2013; Hajar et al., 2014; Park et al., 2013; Robillard et al., 2013), and high activity/impact communications tend to originate from public relations experts who work for those organizations (Harris et al., 2013). Characteristically, these communiques tend to include links to credible health education sites and news sources (Love et al., 2013).
Health-related organizations can also use Twitter to follow and influence policy decisions. For example, Twitter can be used in real time to follow public sentiments relating to pending healthcare legislation (King et al., 2013), and Twitter communiques can be used to influence signature collection campaigns (Yamaguchi et al., 2013). Also of interest to both health-related organizations and professionals is the finding that Twitter altmetrics are consistent with traditional gauges of publication impact such as citation counts (Thelwall et al., 2013).
As individuals, many healthcare professionals use Twitter solely for social versus professional reasons (Hajar et al., 2014; Lulic & Kovic, 2013). Moreover, healthcare professionals do not necessarily embrace Twitter as a communication tool at professional conferences (T. Desai et al., 2012; Nomura et al., 2012). This seems reasonable given that many individuals attend conferences specifically to network face-to-face. In regard to one other small venue use, Twitter has received favorable reviews as a tool to provide private evaluative feedback to residents in clinical settings (B. Desai, 2014).
Discussion
Insights and issues pertaining to Twitter-related research are emerging quickly, and this review represents a time-limited snapshot of what is currently known. In general, there is a great deal of interest in using Twitter to explore health science research questions. Among the investigations that were reviewed, the research approaches were generally in accordance with macro-level pattern recognition (e.g., Kim et al., 2013; Myslín et al., 2013; Nagel et al., 2013; Tavares & Faisal, 2013; Yamaguchi et al., 2013). This is consistent with the type of research that Twitter Inc. is endorsing at major research institutions such as the Massachusetts Institute of Technology, Harvard Medical School, and Boston Children’s Hospital (Bray, 2014).
Also of note is the fact that comparatively few investigators attempted to purposefully establish Twitter accounts, recruit a small group of followers, and use the platform to conduct small-scale investigations (e.g., B. Desai, 2014; T. Desai et al., 2012). This is despite the fact that Twitter account holders can easily download small datasets from their personal feeds at no cost. Reluctance to access these types of small datasets for research purposes probably relates to the fact that Twitter is perceived to be best suited for broadcasting information rather than for interacting with individuals on a more personal level (Brenner, 2014; Neiger et al., 2013).
Struggles involved in working with big datasets were evident among the investigations that were examined. Researchers routinely appeared to limit their samples in accordance with their ability to process and analyze the data. In addition, the challenges of categorizing and interpreting massive amounts of text-based information have not been resolved. In short, although big data is available, the knowledge and skills to make sense of it have not been fully realized (boyd & Crawford, 2012; Vickey et al., 2013).
Another problem associated with the use of Twitter to conduct health-related research is that, despite its geographic reach, users do not represent every sector and echelon of society. Moreover, it is not always possible for researchers to determine what cohort of users they are a tapping into when they access the Twitter database. Geocoding can help fill this information gap; however, knowing the location of users does not always aid in understanding the personal attributes of those individuals.
In general, researchers are encouraged to stay apprised of anecdotal discontent surrounding the use of Twitter communiques for research purposes (e.g., Anderson, 2014; boyd & Crawford, 2012; Mayer-Schonberger & Cukier, 2013). Currently, evidence-based insights pertaining to the ethical use of public tweets for research purposes are lacking, and additional research is recommended (McKee, 2013). For now, researchers are urged to work closely with their Institutional Review Boards regarding the ethical use of Twitter communiques for research purposes.
There is still a great deal to learn about how to use Twitter to improve the delivery of health care and to enhance well-being around the globe. For now, the trend is to use the Twitter platform to conduct big data research. To take advantage of this vast dataset, continued time, effort, and funding will be needed to optimize data retrieval and analysis. In addition, researchers are encouraged to consider ways to conduct small-scale investigations by using investigator-established Twitter accounts.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
