Abstract
The object of this research is to exploit the algorithm of Twitter’s trending topic (TT) and identify the elements capable of guiding public opinion in the Italian panorama. The underlying hypotheses that guide the whole article, confirmed by the research results, concern the existence of (a) a limited number of elements at the base of each popular hashtag with very high viral power and (b) hashtags transversal to the themes detected by the Twitter algorithm that define specific opinion polls. Through computational techniques, it was possible to extract and process data sets from six specific hashtags highlighted by TT. In a first step through social network analysis, we analyzed the hashtag semantic network to identify the hashtags transversal to the six TTs. Subsequently, we selected for each data set the contents with high sharing power and created a “potential opinion leader” index to identify users with influencer characteristics. Finally, a cross section of social actors able to guide public opinion in the Twittersphere emerged from the intersection between potentially influential users and the viral contents.
The digital space is an artificial environment within which human interactions carried by specific algorithms take place. The social action of individuals on the Internet translates into very large (bits) data and data sets within which, however, only a few elements manage to create a resonance effect on the entire digital context. Therefore, as a starting point, this article uses the Twitter trending topic (TT) algorithm, which analyses the dissemination of information within Twitter. In particular, this article aims to determine which are the reference opinion groups in the Italian Twittersphere and, above all, which users have a fundamental role in the general information with viral content capable of influencing the entire network.
This type of analysis has been developed to both thinking contextually to the contents and the social actors. With the aid of computational techniques of data extraction, processing, and analysis, this article focuses on six TTs that help to isolate some fundamental processes and detect the existence of permanent transversal audiences with respect to the different TTs and relevant social actors capable of producing viral content.
The choice to use Twitter is due both to the streamlined and dynamic structure of social media and to the possibility of easily accessing a significant number of public contents. Twitter presents itself as a relationally asymmetric microblogging platform. The diffusion of the contents (i.e., Tweets) must be summarized in 280 characters, and the links between the user profiles are not necessarily two-way but can be nonreciprocal and grouped by friends and followers (Bentivegna, 2014, 2015; Java et al., 2007; Murthy, 2018; van Dijck, 2011). These basic features make Twitter a much slimmer and more flexible social media than Facebook. In fact, Hagan defines it as “a sort of adrenalized Facebook” (van Dijck, 2013, p. 70).
Twitter offers each user a generally public and very concise personal space in the personal description, while the flow of communication is divided into three fundamental actions: (1) production of content (tweet), (2) sharing content from other users (retweet), and (3) production of content by connecting users to specific concepts (mention). Since 2007, a content cataloging system has already been used in the context of the Internet Relay Chat marked with the # symbol and a defined hashtag and whose popularity is conveyed by the TT algorithm. The hashtag is a fundamental part of the syntax of Twitter and has a dual function, one that is evident and the other latent. The first is attributable to the function of the hashtag itself, that is, to make the tweets related to a certain topic identifiable. However, the second is latent and aimed at outlining the audiences that form related specific topics (Bentivegna, 2015). In conclusion, it is possible to affirm that Twitter presents a network of three-dimensional relationships: (1) asymmetric between subjects (friends and followers), (2) asymmetric between subjects and concepts (@mention), and (3) symmetric between concepts and concepts (#hashtag). These three dimensions are referred to collectively as the Twittersphere.
Literature Review
The omnipresence of social structures can be found within the daily behavior of social actors, up to co-occurring in the determination of their own identity (Maretti, 2018; Walther & D’Addario, 2001). In this frame is conceived the figure of the prosumer, the one who simultaneously becomes the producer and user (consumer) of the media product. The prosumer figure finds ample space and is reworked within the digital space (Choi, 2015; Kotler, 2010; Ritzer & Jurgenson, 2008). In fact, it is more difficult to distinguish the difference between producers and consumers in the immaterial and collaborative worlds of Web 2.0, where the hegemony of prosumers is clearer.
The digital prosumer, however, is not a particular figure but defines any user on the network. The action of producing digital content within the web space, in fact, is not limited to the creation of specific online multimedia content (e.g., post, video, or image), but it also concerns simpler and apparently low communicative activities (e.g., a share or retweet, a like or mention or geolocate yourself in a specific place). Even choosing to open one website over another means having also contributed unwittingly to managing it.
In the digital universe of prosumer 2.0 and mass communication (Castells, 2014), a fundamental role is that of the so-called hubs, connectors (Gladwell, 2006), or influencers (Bakshy et al., 2011). Hubs are influential social actors who possess a number of preferential links within affinity networks and are able to influence the rest of the network. In digital social networks, hubs can generate fashion, spread customs, and determine the virality of information.
According to the theories of Barabási (Albert & Barabási, 2002), the action of the hubs within the exchange networks responds to the law of the “rich who becomes ever richer” (i.e., the theory of invariant scale networks; Albert & Barabási, 2002, p. 27). According to this theory, the influencers present within the network will naturally tend to acquire links and, therefore, consent in a greater way than nodes with a lower degree. For this reason, the most popular sites are more visited than the less popular ones and tend to acquire more links than other sites, and blogs with more followers are subject to acquire more contacts than the smaller ones.
In the 1960s, Lazarsfeld et al. (1968) were the first to define the role of specific individuals, called opinion leaders, in influencing public opinion. The peculiarity of these social actors is that their influence is attributable solely to their relational capacity and their exposure to the mass media. The presence of these particular actors redefined and mediated the flow of communication from the medium to the public, which became a two-step flow (Katz, 1957; Katz & Lazarsfeld, 1955). The theory developed by Lazarsfeld (Katz, 1957; Katz & Lazarsfeld, 1955) refers to the world of mass media; otherwise, in the digital space, the role of opinion leaders acquires a different value. In the network, exposure to the means of communication is available to all users, social relations can be mediated by social media, and the communication tools are manifold. Therefore, digital opinion leaders are users who manage to acquire a strategic position in the network to such an extent that their contents become resources capable of attracting the attention of the public and acquiring the power of virality (Choi & Park, 2015; Giurgiu & Barsan, 2008).
The behavior of relevant social actors capable of generating social influence also emerges in the studies on Granovetter’s models of collective behavior (Granovetter & Soong, 1983). According to Granovetter’s threshold model, the emergence of collective movements is based on two conditions: (1) a population of agents with binary choices and (2) the influence of the behavior of others. Therefore, if an individual “instigator” acts autonomously, it will provoke a domino effect such as to generate a macro collective behavior. Opinion leaders have some power over their followers and exercise this power by influencing their behavior (i.e., their choice of action). Therefore, after all the actors have chosen their actions, a decision-making mechanism determines the collective choice, resulting from individual choices but influenced by a relevant social actor (Krassa, 1988; Valente, 1996; van den Brink et al., 2011).
Lazarsfeld’s theory (Katz, 1957; Katz & Lazarsfeld, 1955) becomes the basis for Gladwell’s (2006) studies on the phenomenon of virality. According to Gladwell (2006), there are three factors that lead a phenomenon toward the explosion: (1) the law of the few, (2) the grip factor, and (3) the power of context. According to this theory, the explosion of viral phenomena depends, first of all, on the type of infected people involved, called contaminators. Contaminants must have a large network of connections, be able to create what Merton (1948) and Watzlawick (1984) called self-fulfilling prophecies or be sellers or have a particular charisma that allows them to convince others of their arguments. Opinion leaders must be able to modulate their message in relation to the receiving public and to assess the “fertility” of the context to which the message is being addressed.
Trending Topics (TTs) and Opinion Leaders in the Twittersphere
The polarized structure of the Twittersphere is attributable to both the presence of influencer users and popular hashtags. The popular hashtags are signaled by the TT algorithm, which automatically suggests to the user the topics that are of great interest at a particular time and that will naturally acquire more content by polarizing the topic of discussion (Aiello et al., 2013; Lee et al., 2011). The polarization process also nestles specific users who present the typical characteristics of opinion leaders. In particular, on Twitter, the role of the leader is updated and expanded through continuous interactions of digital social actors (Murthy, 2012, 2018) that build and maintain around the leader a relational network constantly fed via three main tools: follow, retweet, and mention (Bentivegna, 2014, 2015; Cha et al., 2012).
The Research Project
The exposed theoretical framework outlines a series of ontological and phenomenological issues, concerning the development of public opinion in the digital space. First, it is necessary to ask whether the Lazarsfeld paradigm (Katz & Lazarsfeld, 1955) of the two-stage flow of communication still represents a valid point of reference for the social dimension. Second, it is necessary to evaluate the characteristics of opinion leaders (or influencers) within the same structures. Finally, it is necessary to ask what the typical language is of the public within the specific affinity networks, without forgetting that each digital structure present on the web has its own public and informational context and carries a specific form of communication (Rainie & Wellman, 2012). The objective of this contribution is to make operative the issues previously exposed in the Italian Twittersphere. The research design translates into the following research questions and hypotheses:
Data Collection
During the period from November 2018 to February 2019, a data set containing 323,197 tweets was collected through a streaming application programming interface (API) process in relation to a selection of hashtags suggested by the TT algorithm. The hashtags concern a heterogeneous series of news and current events relating to the Italian political, cultural, and social context such as #Battisti: regarding the arrest and repatriation of the terrorist Cesare Battisti; #Orango: concerning the debate on the sentence of 18 months’ imprisonment of Roberto Calderoli, accused of having defined Cécile Kyenge with the name Orango (Orangutan). #Libero: this hashtag was taken over by a TT on January 23, 2019, following the appearance on the first page of the newspaper, Libero, with the headline “Cala il Pil, aumentano i gay” (The gross domestic product falls, gays increase); #Rousseau: regarding the voting phase in the Rousseau platform and the decision to intervene or otherwise against the minister Matteo Salvini in the Diciotti case; # Sanremo2019: relating to the San Remo 2019 festival final and the controversy over the results; and # SeaWatch3: including all the opinions of Twitter users regarding the blockade against the Sea Watch 3 ship.
The process of data mining Twitter was carried out through the use of the online platform, Social Grabber, through which it was possible to download data in streaming modality and in relation to specific key words identified by language and time segments. The download phase made it possible to set up six data sets, one for each TT. Tools programmed in the Python environment were exploited to derive (a) the hashtag adjacency matrix for each TT, (b) the overall hashtag adjacency matrix for all the six TTs, (c) an electronic spreadsheet containing all the tweets with the relative users and the number of retweets and retweets with text, and (d) a data set containing the list of all the users who produced content, the number of posts produced on the TT in the observation period, the number of followers and of friends, the number of times that its content was retweeted, and the number of times it received a mention.
The Research Stages
A specific analysis technique was used for each research segment.
An analysis of the six semantic networks of each sample aimed at identifying the hashtags that co-occurred most frequently with the main TT with the aid of the Gephi software (Bastian et al., 2009). Finally, the transversality of certain specific elements was detected and a semantic graph was constructed to illustrate the hashtag network and to identify, through centrality measurements, the influential elements present together in the TTs.
The reconstruction of the relational structure of the various topics was carried out by considering the total number of elements (consisting of tweets and retweets) present within each topic in relation to the number of individual tweets. Finally, for each tweet, the retweet frequency was evaluated on a scale: 0−10, 11−100, 101−200, 201−300, and from 300 upward.
In this phase of the research, the objective was to identify the elements capable of delineating the role of the opinion leaders within the data sets. The reference literature (Cha et al., 2010; Choi, 2015; Leavitt et al., 2009) has not been uniform in its evaluation of the indicators that define opinion leaders. In fact, the follower numbers alone cannot be considered as a value that is able to indicate conclusively the role of the influencer because there are elements that have a certain number of followers exclusively for their social role, so they do not have an interactive audience. It is no coincidence that authors such as Cha (Cha et al., 2010) and Avnit (2009) have talked about the million follower fallacy. Cha (Cha et al., 2010) defines effectiveness in terms of indegree influence, retweet influence, and mention influence. Indegree influence measures the size of the audience for each user, retweet influence indicates the ability of each user to generate viral content, and mention influence implies the ability of every user to be involved in other conversations (Cha et al., 2010).
The indicators defined by Cha (Cha et al., 2010) are reference points for identifying potential opinion leaders in the Twittersphere. However, Lazarsfeld’s general theory (Katz & Lazarsfeld, 1955) defines opinion leaders as the first consumers of media products. Furthermore, as Gladwell (2006) stated, the explosion of a current of thought is determined not only by the law of the few but also by the grip factor and the power of the context.
In relation to these three paradigms (Cha et al., 2010; Gladwell, 2006; Katz & Lazarsfeld, 1955), it was decided also to consider other indicators, creating a synthetic index of “potential opinion leader.” These were the number of post products (media consumption), the number of followers, the number of times posts have been retweeted, and the number of mentions received within the data set (the law of the few). The index is obtained by normalizing the indicators, rescaling them between 0 and 1. The standardization process was realized with the min–max elementary indicators: Z = (x − min(x))/(max(x) − min(x)), and the single elementary indicators were aggregated using the arithmetic mean.
Subsequently, the results of each indicator were cross-referenced with the value of the most retweeted contents (the power of the context), and the common elements present in both data sets (the grip factor) were considered as influencers of the one specific TT and classified in relation to their social role in the digital space of Twitter.
For each opinion leader identified in Phase 3, the specific scope of influence was detected through the qualitative study of their public profile. Moreover, through the analysis of the content of its production in the Twittersphere, the type of opinion expressed in the public sphere of discussion was identified and classified.
Results
The Co-Occurrence of Hashtags
As illustrated in the previous paragraph, the first research step analyzed the hashtags that most frequently co-occurred within the six TTs identified: (1) #Battisti, (2) #Orango, (3) #Rousseau, (4) #Sanremo2019, (5) #Libero, and (6) #SeaWatch3 (see Figure 1). The co-occurrence analysis—using the Gephi software package—illustrates a semantic scene around each network composed of certain hashtags that redefine the topic being discussed, and others that are more general, some of which can be traced back to broader topics of discussion.

Six semantic graphs, one for each trending topic (#Libero, #Orango, #Battisti, #SeaWatch3, #Sanremo2019, and #Rousseau) visualized by Yifan Hu Proportional layout.
The results also present six dimensions that are not completely separate; thus, some hashtags are common to more than one data set with a preferential position in each network in terms of weighted degree, closeness centrality, and betweenness centrality 1 (see Table 1). The hashtag #facciamorete is one of the most used in all six TTs. It was created in December 2013 by the digital activist user, Marco Skino (@MPSkino), who in his personal page describes it as a network son of a cultural project aimed at creating a community of commentators who share an anti-fascist philosophy and who are in favor of protecting the Italian Constitution. Another example of transverse hashtags is #salvini, which is the most used in all six data sets. However, it is necessary to point out that Matteo Salvini (@matteosalvini) has a decisive role in four of the six events that led to the birth of the TTs, including, as Secretary of the Lega Party (to which Roberto Calderoli also belongs), the affair related to the hashtag #Orango.
Analysis of the Most Frequent Terms for Each Trending Topic.
To show the transversality of certain elements, an adjacency matrix was built between all the hashtags present within the six TTs, and this was displayed in a graph. From the analysis of the network of all the hashtags present within TT (excluding replication), an indirect small-world graph emerged, composed of 1,735 nodes and 36,933 links, with a diameter of four degrees of separation and with a clustering coefficient of 0.587. The semantic network is composed of semantic links that can be defined as strong links because they are inherent in the same thematic area and in the weaker links of more generic content. Within the graph, eight hashtags have been selected that show a preferential position within the network, both in terms of number of bonds and in terms of weighted degree and betweenness centrality (Table 2).
It should also be noted that two hashtags, #facciamorete and #salvini, appear both as more frequent elements within all six TTs and between nodes that have a preferential position within the hashtag graph. We can, therefore, deduce that #facciamorete and #salvini are two hashtags of relevance, both for structural issues and for the type of content. In fact, their coexistence in several areas of conversation as a frequent element and their placement in a preferential position with respect to the relationship network define these hashtags as elements of influence in the dissemination of information. To the quantitative data is to be added a qualitative reflection concerning the reference public. In fact, #facciamorete and #salvini, by a precise declaration of intent, create two opposite poles of discussion and opinions and are placed as bearers of two specific opinion subculture spheres.
Nodes With a Greater Value of Degree and Betweenness Centrality.
The Relational Structure Inside a Hashtag
The conversational structure of the six TTs was revealed in the relationship between the total number of data of each data set, the number of tweets, and the number of retweets. The frequency of sharing was analyzed for each retweet. From the relationship between the number of tweets and that of retweets, the number of tweets (i.e., the content at “first hand”) had an average of 20.47% compared with 79.53% of retweets (Table 3).
Number of Tweets and Retweets (Internal and Percentage) for Each Trending Topic and the Average Value.
Furthermore, the total number of tweets that had a number of particularly relevant shares represented a very small portion of the whole data set. In fact, on average, the content was retweeted more than 300 times (0.17%); therefore, it reached a wider audience, while 94.57% of the content tended to have a niche visibility not greater than 10 social actors each (Table 4).
Retweet Frequency (Internal and Percentage) for Each Trending Topic and the Average Value.
Therefore, it is possible to conclude that despite the breadth of the data sets (323,197 posts), there is little first-hand content compared with the value of the shares. Moreover, few of those contents are able to reach the wider public and, thus, can be considered significant for determining spheres of influence in the Twittersphere. Finally, it is necessary to point out that the values classified as greater than (>) 300, despite being few in number, reach very high sharing values (Table 4).
In view of the highlighted results, it can be stated that the conversational structure of all six TTs represents (Figure 2) a free-scale network 2 (Albert & Barabási, 2002). This is composed of a few nodes that have many links representing preferential elements. These have more likelihood of generating greater influence in the network than many nodes with few links, which will have a low level of influence.

Occurrences curves of the data elaborated in Table 4 in percentage value.
The Opinion Leaders
As previously stated, it is not easy to define useful indicators to identify the figure of the opinion leader within a data set. Indeed, taking the follower number exclusively as a point of reference is completely reductive. Several empirical studies have elaborated synthetic indices using follower, retweet, and mention (Cha et al., 2009; Choi, 2015; Leavitt et al., 2009; Marchetti, cited in Bentivegna, 2014). In the present study, and in accordance with the paradigms of Gladwell (2006) and Katz and Lazarsfeld (1955), it was decided not only to select indicators capable of describing the absolute role of the user but also to evaluate media production, the grip factor, and the power of the context. In other words, it was not only the potential of the users as opinion leaders that was evaluated but also their role as influential elements in a specific context. In this regard, a summary index was drawn up, consisting of the following: the number of followers, the number of posts in the data set, the number of mentions, and the value of the retweets of content relating to the TT (Figure 3).

A synthesis schema illustrating the process of identifying the opinion leader within the six trending topics.
The index revealed the number of users with the characteristics of the opinion leader as well as the content actually quoted by the public that reached a retweet number greater than 300 (the grip factor). A sample of users was selected from both lists. The analysis carried out highlighted a sample of 26 users, representing the influential elements relating to specific TTs.
In this regard, three observations of an empirical nature can be made. The first concerns the transversality of some actors who are influential in many discussions (i.e., @matteosalvini, @Dio, @lauraboldrini, @ficarraepicone, @francescatotolo, and @vitocontesi). Transversality also encompasses other users (i.e., @makkok, @MPSkino, @Iperbole, and @Manginobrioches) who, despite being in these specific contexts, did not produce content with great potential for virality. The second observation concerns the heterogeneity of the sample; the selected influencers have different social roles: politician, journalist, cartoonist, activist, singer, show character, and comedian. The third observation concerns the users @matteosavinimi and @MPSkino, who, besides being influencers, also represent the central node of the two hashtags #salvini and #Facciamorete, both of which were preferential elements in the hashtag semantic network (Table 5).
Identified Opinion Leaders Referenced by Trending Topic and Their Classification.
The Social Role of Opinion Leaders in the Digital Space
In the introduction, we reflected on the role of opinion leaders and the kind of influence they can have in defining specific spheres of opinion and collective actions (Granovetter & Soong, 1983). Gladwell’s (2006) law of the few refers to certain subjects with specific qualities who are capable of creating fashions or influencing others in particular ways. Bentivegna (2015) underlined the fact that the digital space, in particular that of Twitter, allows each Internet user the possibility of becoming an interpreter of a specific public sphere. Finally, Lazarsfeld (Katz & Lazarsfeld, 1955) defined the opinion leader as a molecular leader because they have no well-defined social characteristics and can belong to different socioeconomic strata.
In the final phase of the research, the digital role of opinion leaders was considered by analyzing the Twitter profiles of identified influencers and classifying them in relation to three macro categories: politicians, activists, and experts (Table 5). In the category of politicians, there are users who exercise a role within political movements and parties, including those who possess and exercise an institutional role. Specifically, this area has two polarities: the right wing, as represented by @matteosalvinimi, and the left wing, as represented by @AlessiaMorani, @lauraboldini, and @ckyenge.
The activist category is the widest and most heterogeneous, and it includes all the users who—while not exercising a specific role—manifest the opinions of populist right-wing politicians and progressive leftists or express themselves using satire. Activists include both ordinary citizens and public figures. Specifically, we distinguish @intuslegens, @benq_antonio, @francescatotolo, and @vitocontesi as populist right-wing activists and whose content is also in line with the principles of the populist right. Although belonging to different categories, @vitocontesi’s content is linked to that of @matteosalvinimi. On the other side, progressive activists include both public figures (including the cartoonist @makkox, comedians @ficarraepicone, and public figure @vladiluxuria) and citizen activists (e.g., @SarahConnor_18, @CaputaLuca, @ giuliaselvaggi2, @MPSkino, @Iperbole, and @manginobrioches). Furthermore, there is the subcategory of political satire, which includes users (e.g., @Dio and @frasidiOsho) who express political opinions, in this case of a progressive nature, through satire. For instance, @frasidiOsho is actually a very active portal (also on Facebook) that produces content mainly through the communication tool of the meme.
Finally, experts are users who declare a cultural background that makes them experts in certain topics of discussion. In this category, we included the journalists @mattiafeltri, @alessiarotta, @marcellofoa, @alessioviola, and @CucchiRiccardo as commentators and opinion makers, respectively, in the hashtags #battisti, #rousseau, #Libero, and #SanRemo2019. We also included the singer @frafacchinetti as an active commentator in the hashtag #Sanremo2019.
In terms of the qualitative analysis of the profiles of the identified opinion leaders, it is necessary to underline that their real identity is sometimes unknown. Anonymous activists like @intuslegens, @Dio, @MPSkino, and @manginobrioches do not show their image and offer no personal description or geographical location.
Conclusions
Some of the fundamental elements that regulate the Twittersphere have emerged from this study. The structure of the analyzed discussion networks presents—in its double semantic dimension—the same composition dominated by a few elements with strong relational power. The study, therefore, allows some conclusions to be drawn as to how dialogue is established and maintained in the Twittersphere. The first consideration is the applicability of the Katz and Lazarsfeld (1955) paradigm on the two-stage communication flow and of its use in the digital space (Choi & Park, 2015). Empirical evidence has revealed the presence of molecular social actors who play the role of a sounding board and polarizers of public opinion by generating networks with well-defined contours. Opinion leaders cannot be regarded exclusively as carriers and interpreters of opinions and trends. They are central and aggregating nodes of specific subcultures of opinion and play a predominant role in the management of information and disinformation (Boccia Artieri, 2012).
Having described the real presence and influence of these users, it is necessary to underline the following point: To evaluate the virality of content, one cannot refer solely to the role of the opinion leader; the actual activity of users in a particular area of discussion and the credibility granted to them by the specific audience must be taken into consideration. It is necessary to think not in terms of a receiving public but of receiving publics. Each web medium is the bearer of a precise type of communication, and within it, there are several different audiences. Therefore, the role of the opinion leader is not absolute; it acquires value when it is legitimized by a specific segment of opinion (Gladwell, 2006).
The last but not least consideration concerns the part played by the hashtag in the Twittersphere. According to Bentivegna’s (2015) theory, the hashtag has a dual function: one that is evident and one that is latent. However, if the latent function involves an ad hoc interpretative frame, it is then necessary to add specific niches of discussion that, although defined by the hashtag, are not extemporaneous elements of context but distinctive and transversal to specific affinity networks. Especially in cases such as #salvini and #facciamorete, it is too simplistic to think in terms of information containers or an ad hoc public; rather, they should be visualized as cultural identifiers of certain publics in particular contexts of opinion. Another relevant aspect is that of the structural similarity that exists between the semantic network of hashtags and the structure of the relationship between users. This reflects the model of the invariance scale networks defined by Albert and Barabási (2002), which continues to be the universal, distinctive, and confirmatory element of all digital organisms.
Finally, the research results show that @matteosavinimi and @MPSkino play a particularly influential role in the Twittersphere. The two users appear simultaneously as opinion leaders of some TTs and references of the two transversal hashtags, #salvini and #facciamorete, which in fact describe two opposite opinion polls. For this reason, it is necessary, in future applications, to deepen the role of these two polls as possible permanent communities of opinion in the Twittersphere and reference point to analyze specific collective mass behaviors (Granovetter & Soong, 1983) that arise between the digital space and the offline dimension.
Footnotes
Data Availability
The adjacency matrices relating to the analysis of the co-occurrences of the hashtags are available at
. The data set extracted directly from Twitter in streaming application programming interface system in JavaScript Object Notation format is kept by the authors. For a possible replication of the research, you can contact them at the following addresses:
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Software Information
The tweets’ data set was collected using Socialgrabber, an online service platform that provides a user-friendly graphical user interface to use the publicly available Twitter Streaming application programming interface (API). The platform allows to collect data scheduling data acquisition jobs and setting filtering parameters, and it allows to export them in several formats. We exported that in “JavaScript Object Notation Lines” format.
3
Twitter provides for each tweet several fields (http://jsonlines.org/
with data referring to the user, the text, the time, and the entities occurring in each tweet (e.g., hashtags, mentions, and URLs). The Twitter Streaming API, in addition to just normal tweets, also provides quotations and retweets as a data entity similar to normal tweets. A retweet is when the user republishes a post that another Twitter user has written. It is a way of amplifying the signal, so more people hear the original message. A quotation tweet is a kind of retweet. While a simple retweet merely shares another person’s tweet, a quotation tweet lets you share another person’s tweet and add your own comments. Quote tweets are sometimes referred to as a “retweet with comment.” In both cases, Twitter Streaming API, in addition to the mentioned tweet data, provides additional fields conveying information about the original tweets. For our analyses, the retweet and the quotation counts for each pure (or original) tweet were computed, summarizing the number of retweets and quotations, respectively. However, the additional text in quote tweets was kept in the data set as the user’s custom text part is like a normal tweet. In some cases, retweets and quote tweets refer to pure tweets that were not grabbed in the streaming collection process because either the pure tweet was posted before the stream collection started or it does not include any of the keywords used in the filtering. Since the original tweet entity information is present in the retweets or quote tweets, the script allows to extract and add those pure tweets to the data set. The script allows to extract the tweet information and to perform descriptive statistics such as the number of posts sent by each user, the number of mentions, retweets, and quotations referred for each user. As data are gathered in a streaming way, the stats about each user (e.g., the number of followers and followed people) may change between tweets. So, for each user, only the value present on the most recent tweet is used for these statistics. The hashtags co-occurrence matrix is created, merging without repetition all the different data sets and extracting the hashtag list present in the dedicated field of each tweet.
4
