Abstract
Social media platforms, such as Twitter (now X), are a major source of communication. Identifying communicative intentions is useful, as it encapsulates the latent motivations that drive text creation. This intention is also helpful in understanding the message, context, and audience. This study proposes a method for detecting communicative intentions in tweets using Jakobson’s language functions. We constructed a meticulously annotated dataset, drawing from the extensive RepLab2013 corpus. Our dataset underwent rigorous scrutiny by linguistic annotators who analyzed over 12,000 tweets individually. These experts identified the dominant language function within each tweet by employing diverse strategies to ensure precise labeling quality. The outcome demonstrated a noteworthy Kappa agreement score of 0.6, reflecting a strong inter-annotator reliability. Subsequently, these functions were mapped to the corresponding intention categories. We employed logistic regression and support vector machines (SVM) algorithms to classify intention in tweets and explored various pre-processing techniques, incorporating n-grams and bag-of-words representations. Furthermore, we expanded our research using pre-trained large language models, incorporating the latest state-of-the-art techniques in natural language processing.
Keywords
Introduction
The communicative intention is the purpose that a participant in a communicative act aims to achieve through a discourse. The intention shapes the addresser’s discourse, as their linguistic acts will be directed towards achieving the intended purpose (even unconsciously) while influencing the addressee’s interpretation [3]. This intention can, in principle, be detected automatically through analysis using natural language processing of texts, such as social media discourse [15]. Social media platforms, such as Twitter, have become a major source of communication in modern society. As people use these platforms to share their thoughts, opinions, and experiences, it is important to understand the underlying motives behind the text being produced. In [13] the authors defined the relations between six factors and their corresponding communicative functions that allow effective verbal communication. The relations highlight the importance of each factor taken in a language function as follows: (addresser, emotive), (addressee, conative), (contact, phatic), (message, poetic), (context, referential) and (code, metalingual). In general, texts serve multiple language functions and have various communication intentions. However, one function tends to predominate [2]. According to Arcand [2], if the main function in a text is referential, then the intention is informative. If the main function is emotive, the intention is to express the addresser’s state. If the function is conative, the intention is to influence; if the function is phatic, the intention is to engage; if the function is poetic, the intention is to stylize the message; if the function is metalingual, the intention is to clarify. To create our dataset, we analyzed 12,653 tweets in Spanish from RepLab2013, where linguist annotators determined each text’s most dominant language function.
Our primary objective was to create a prediction model based on supervised learning to automatically classify communicative intentions in tweets. Our principal contributions are the following: A dataset based on RepLab2013 labeled with the language intention category. Classification models that identify communicative intention using traditional machine learning and deep learning methods.
Detecting communicative intentions by studying the language functions in tweets has significant applications for businesses, organizations, and individuals who use social media platforms as tools to understand the purpose of a tweet in reputation management and marketing contexts. Another application is to help understand the audience’s intentions, improve engagement, and adapt content strategies.
The remainder of this paper is organized as follows: Section 2 presents related work on intention and language function detection using machine learning and natural language processing in social media; Section 3 details Jakobson’s language functions and the intentions related to these functions. We also show the details of the dataset creation process, including the agreement between annotator values. Section 4 describes different experiments for obtaining our classification models, including tweet pre-processing methods, text representation models, and classification algorithms. Finally, Section 5 presents our conclusions and outlines potential directions for future research.
Few studies have addressed the classification of intention or language functions in social networks. Martis and Alfaro [18] proposed eight tweet intention types: news reports, news opinions, publicity, general opinions, shared locations or events, chat messages, questions, and personal messages. However, these are more related to the user’s intention in their use of this social network. The dataset used consisted of 5,222 Spanish-language messages provided by an Analytics business and labeled manually. They built two models using the Naive Bayes and SVM algorithms in WEKA. The performance difference between the models was insignificant, with 98% of precision measure in Shared Locations and the worst precision measure in General Opinion classification at 73.5% Karpov et al. [15] created a dataset for their study and used VKontakte, a popular Russian social network, as the source of discussions and text. Raw data were downloaded using VKMiner, and rigorous filtering was performed to remove discussions with less than 40 phrases and meaningless comments, such as empty messages or photos, instead of texts. Their annotation consisted of two labels for each message: intention type and direction of intention (referring to the owner of the intention). They used the intention categories proposed by Oleshkov et al. [14], typifying a communicative strategy as a general goal of a speech act with additions [19] having four categories for direction, 25 intention types, and five intention super-category classes. They built two models using LSTM and CNN deep learning architectures, obtaining better results with LSTM and super-category classes, achieving a precision of 35%.
Ismaeil et al. [12] revisited the language functions, expanding the original definitions and using Politics subreddit 1 , a popular forum for political news from the U.S.. They restricted attention to 165,000 short comments consisting of at most two syntactic phrases. e.g., “I see your point” (NP-VP). Using three pre-trained annotators, their final dataset comprised 4,482 distinct utterances, which they called messages. The average precision was 83% using a graph-based, semi-supervised label propagation framework with a Modified Adsorption (MAD) algorithm.
Related work has dealt with different social networks, text filtering criteria, labeling methods, in-text pre-processing, classification algorithms, machine learning approaches, categories, and conceptual models to identify intent automatically. Consequently, this is reflected in the disparity of results in the literature.
Corpus of communicative intention in Spanish
The RepLab2013 dataset contains tweets in English and Spanish and was originally used to evaluate online Reputation Management Systems [1]. It consists of tweets related to entities, such as companies, organizations, and celebrities on Twitter. The selection was made to offer various scenarios for reputation studies. The creators of RepLab2013 were tasked with sifting through the tweet stream for potential mentions of the entity, filtering those that referred to it, and then clustering the remaining tweets by topic to rank them based on their potential as reputation alerts. Crawling was performed from 1 June 2012 to 31 Dec 2012, using canonical names as queries for each entity.
In this research, we take a subset of the RepLab2013 corpus with a total of 12653 tweets, selecting a set of Spanish tweets, including four domains: automotive, banking, universities, and music.
Annotation methodology
Throughout the process, we confronted the inherent complexity of language function labeling, as certain cases presented multiple interpretations. This complexity was largely influenced by the cultural context embedded within the tweets and the annotators’ individual perspectives. Although annotators can assign more than one language function in the tagging process, the final version of the corpus consists of a single label, determined using the methodology described in [21] as a reference.
Elements in communication and functions of language
Figure 1 schematizes the factors involved in verbal communication according to Jakobson. He distinguished six elements in any speech event or act of verbal communication. An addresser sends a message to an addressee. To be operative, the message requires a context referred to, approachable by the addressee, and either verbal or capable of being verbalized; a code fully, or at least partially, common to the addresser and addressee; and finally, a contact, physical channel, and psychological connection between the addresser and addressee enabling both of them to enter and stay in communication [13].
Referential: is typically associated with conveying factual information or referring to the reality outside the language system. In this function, the primary aim is to represent objects, ideas, or states of affairs in a way that the listener or reader can understand [17]. Examples of referential messages include observations, opinions, and factual information.
Poetic: it focuses on the aesthetic and artistic aspects of communication. In this function, language is used not just to convey information but to create a particular effect or emotional response. Attention is given to the form and structure of the words and phrases, including elements like rhythm, rhyme, alliteration, and metaphor [17]. This includes double meanings, verses, proverbs, and sayings.
Emotive: centers on the speaker’s feelings, judgments, or emotional state. In this function, language is a tool for expressing subjective experiences and emotional reactions. The aim is to evoke empathy, understanding, or a particular emotional response from the listener or reader [10]. This function is crucial in personal communication, storytelling, and any context where emotional expression is important.
Conative: is oriented toward prompting action or a response from the listener or reader. This function primarily focuses on the “you,” or the second person, as it aims to persuade, command, request, or question. Messages are conative if they represent orders, demands, advice, or wishes.
Phatic: it establishes, maintains, or closes a communication channel between the sender and receiver. It doesn’t aim to convey information or express emotions per se but rather focuses on the social aspect of communication. These messages primarily serve to establish, prolong, or discontinue communication, to check whether the channel works, to attract the interlocutor’s attention, or to confirm his/her continued attention.
Metalingual: it involves using language to talk about language itself. The metalingual function allows both the sender and receiver to ensure they are linguistically on the same page, facilitating clearer and more effective communication. Examples include asking for definitions (“What does ‘ambiguity’ mean?”), explaining grammatical rules (“The word ‘running’ is a gerund in this sentence”), or clarifying the meaning of an utterance (“When I said ‘that’s interesting,’ I meant it was surprising”).
Predominance of Jakobson functions
We used language functions of Roman Jakobson’s to identify the intentions of a message. In [2], the authors mentioned that “the dominant function is the one that answers the question of what intention was this message transmitted? and [...] The secondary functions are there to support it.” The intention associated with each fragment must be distinguished from the overall intention, which is “a sentence or a series of sentences that correspond to an intention”. [11].
Table 1 presents the mapping (as defined in [2]) of the language functions in communicative intentions labels as follows: referential - inform, conative - influence, emotive - emotional, poetic - stylish, phatic - engage, metalingual - clarify. We included a Spanish tweet example for each language function. In our labeling process, we deliberately incorporate the use of irony and sarcasm as linguistic resources. We recognize their significant role in adding stylistic nuances to messages, which is in line with the principles of stylish intention corresponding to poetic function as defined by Jakobson.

Functions of Language, each focuses on different communication elements.
Below, we detail the six language functions that Jakobson identified for understanding the nature of human communication.
Intention in Language Functions
In the initial phases before the labeling process, we evaluated the competency of each annotator. This involved the creation of a detailed annotation guideline and evaluating a subset of 75 tweets to ensure the understanding of the guideline. During the previous evaluation, we observed that different annotators failed to assign the same language function to the same tweet, which complicated the identification of a singular predominant language function within certain tweets.
The group of annotators included six linguistics scholars affiliated with an expert from Grupo de Ingenierıa Lingüıstica at the Universidad Nacional Autónoma de México. Figure 2 shows the process of creating the subset of 75 labeled tweets, which was used later as the control set during the annotation process.

Annotation Methodology. Six annotators. Every document had three labels from the three different annotators.
For the annotation process, we used the Toloka platform
2
. When the annotation finished, every tweet had three labels corresponding to the classification of the three annotators. To identify a single language function for each tweet, we designed these criteria (based in [21]): When all annotators assign the same label to the tweet, this label is used as the gold standard. If there is a label that has the majority of votes, then this label is used as the gold standard. If a majority vote cannot be reached because all annotators tagged different language functions, then the label with the highest priority is selected as the gold standard. When at least two annotators identify multi-labels for a single tweet, then the label is determined by ranking the language function labels according to popularity among the annotators. Then the label with the highest priority is selected as the gold standard. Disagreements about annotation types were resolved using the highest priority language function label.
We defined the priority as directly proportional of the frecuency of the language function, and we modified rules 3 and 4 to have a single label for each tweet and obtained the best ranking performance based on the F1-score macro. This methodology is summarized in the rules in Table 2.
Examples illustrating the algorithm by which consensus among individual annotations has been achieved to form a gold standard corpus
To evaluate the labeling quality we calculated the inter-annotator agreement. For this, we divided 12653 tweets among six annotators. Since there were a total of six annotators, and the assignment of tweets was done randomly, to obtain the inter-annotator agreement, we chose the group of three annotators who labeled the highest number of tweets in common. Then we filtered the tweets with only one label. This process results in 2540 tweets representing 20% of the total tweets. Finally, we computed the Cohen Kappa agreement, which resulted in a substantial agreement because is in the range (0.61 <x <0.8) defined in [25], as shown in Table 3 all pairs of annotators in the study have shown “substantial agreement” in their annotations. This implies that there is notable consistency in the way these annotators have categorized the language functions in tweets, suggesting that the annotations are reliable and there is significant alignment in their criteria or interpretations.
Cohen Kappa measure
Cohen Kappa measure
Table 4 presents the general statistics for the different language functions in the tweets. Regarding the tweets count, the referential function had the highest number of tweets with 7745, followed by the emotive function with 2654 tweets. By contrast, the metalingual function had the lowest number of tweets, with only 29. With respect to the average number of words, the referential and conative functions have the highest average per tweet, with approximately 18.16 and 17.87 words, respectively. The phatic function has the lowest average number of words per tweet, with around 15.05 words.
General statistics computed from word counts on each tweet
General statistics computed from word counts on each tweet
The emotive function had the highest standard deviation of approximately 6.26, indicating greater variability in the number of words used in tweets for this function. The referential and phatic functions also have relatively high standard deviations compared to the others.
These statistical results reveal different tweets across different language functions; however, the number of tweets labeled in each language function can also influence these statistics.
To generate an automatic classification model for the automatic detection of communicative intention in tweets using the language functions labeled in the corpus based on Supervised Learning, we performed the following experiments: We conducted a comparison of different common pre-processing methods, using Logistic Regression and SVM to generate classification models. We used the macro F1-Score as the performance metric and performed a T-test to compare dependent population means. The reference point was the method without any pre-processing steps, with the alternative hypothesis being that the means are different. Subsequently, we conducted the same statistical analysis at the classification algorithm level. Using the pre-processing method and algorithm that yielded the highest F1-Score value in the previous step and which also supported the hypothesis of differences in means, we conducted a comparison of different feature sets to characterize the tweets. These feature sets included words, characters, and Part-Of-Speech, n-grams with different n-gram ranges and minimum document frequency. Similar to the previous step, we performed a statistical analysis to compare dependent population means with the aim of establishing the significance of the mentioned parameters. Taking the parameters from the two highest F1-Score values for each feature set, which also supported the hypothesis of differences in means, we compared the unions of combinations of different feature sets using the aforementioned parameters. We conducted the same statistical analysis as in the previous steps with the aim of establishing the significance of these feature set combinations. Finally, we also compare the performance of different pre-trained large language models for automatic intention classification.
In the domain of automated document classification, we can find common pre-processing techniques that encompass steps such as tokenization, removal of special characters, among others [6, 20, 23, 24].
To discern the optimal pre-processing approach for our dataset, we conducted a series of experiments to evaluate the efficacy of diverse text pre-processing methods employing the following steps: link, hashtag, username, stop, minus, mentions, emoticons, emojis, and symbols defined as follows: Symbols: We replaced each symbol by its description in Spanish. We consider symbols as punctuation marks that can help identify language functions. Emoticons: We replaced the emoticons with the corresponding text. In the case of having defined an emoji, we used the text of the emoji in Spanish; otherwise, we used the value translated to Spanish from the table defined in
3
using a dictionary of emoticons defined in the list of emoticons. We used the emoticon’s textual meaning instead of the polarity word used in sentiment analysis [8, 22]. The idea of improving classification performance comes from the experience of adding emojis text and obtaining good results for gender classification [24], where we used emojis as text. However, because the tweets in the dataset used were collected in 2012, and Twitter enabled the use of emojis until 2014 there were no emojis in the tweets. Example: Tweet: Para Minus: We changed all capital letters in the text to lowercase text. Link: We replaced urls with the placeholder “link”. Example: Tweet: Para Hashtag: Analogous to the link, the hashtags contained in the text are replaced by the placeholder “hashtag”. Stop: We removed stopwords in the tweet based on the same list as the article [24]; a stopword is a word that is considered common and has no significant semantic value in the language analysis task. Examples include: el, la, de, en, a, y, and o (in Spanish). Username: Similar to the link and hashtag, user names are replaced by the placeholder “usermention”.
Through a comprehensive exploration involving the analysis of 128 combinations of preprocessing methods to build classification models using unigrams with classification algorithms Logistic Regression (LR) and Support Vector Machine (SVM) classification algorithms under a 10-fold cross-validation. We employ the F1-Score Macro metric because we are interested in assessing the classification performance for all classes, considering the class imbalance.
In Table 5, a summary of the individual pre-processing step results is presented, highlighting the combinations that achieved the highest average F1-Score values in bold, and finally, the combinations that obtained the lowest average F1-Score values. We conducted a statistical analysis using all combinations and a paired-sample T-test with a significance level of α = 0.05. The alternative hypothesis considered differences in means, with the reference being the method that did not incorporate any preprocessing steps. As we can observe in Table 5, the results in bold represent the two preprocessing methods, symbols and emoticons and symbols, emoticons, and username, which achieved the highest values compared to the reference for each algorithm, and their means are significantly different. On the other hand, the methods minus, link, hashtag, and stop and link, hashtag, and stop yielded the lowest values. Additionally, it is worth noting that preprocessing that only involved converting the text to lowercase resulted in the same average value as the reference.
Macro F1-Score Average using 10-fold cross-validation with bag-of-word features and different pre-processing step combinations
We consider this phenomenon because the pre-processing step for emoticons enriches the classifier with more information instead of removing symbols from the text to build the term-document matrix. Regarding the symbol pre-processing step, one of the ways in which we can identify the intention of influence is through statements; in this way, exclamation marks serve to distinguish it. On the other hand, the intention to engage typically involves questions characterized by question marks. Regarding the hashtag and stopword pre-processing steps, we consider that replacing the content of hashtags with the term “hashtag” introduces noise into the classification process. This arises from all hashtags being assigned the same uniform value, consequently reducing the informative content for classification. Ultimately, the removal of stopwords generates noise in the stylistic and clarifying intents, as they exhibit a high frequency of functional words. This can be seen in the confusion matrices of the classification models shown in Fig. 3. Using emoticons, minuses, usernames, and symbols, in addition to being able to classify the engage and clarify intentions, performs better for emotional, influence, and stylish intentions.
Subsequently, we conducted a statistical analysis to compare the performance of the classification algorithms, using the average of cross-validation results for each combination and algorithm. This analysis resulted in a p-value of 9.317e-62. Consequently, we can conclude that the results are statistically different. Due to the fact that Logistic Regression achieved better average results than Support Vector Machine (SVM), we have selected Logistic Regression for the upcoming experiments.

Confusion Matrix of the highest and lowest F1-Score Macro metric.
After defining the best pre-processing, which includes {symbols, emoticons, and username}, and selecting Logistic Regression as the classification algorithm, we used word n-grams, character n-grams, and POS tags n-grams as features for the Logistic Regression. We conducted experiments encompassing different n-gram ranges and defined a threshold for the minimum word frequency in the document (min_df) to assess the performance in F1-Score Macro for each feature set. For the statistical analyses of each feature set, we took as reference the lowest average value from the smallest n-gram range.
As shown in Table 6, both unigrams and n-grams (with n varying from 1 to 2) achieved significantly higher values compared to the reference, which in this case consists of unigrams with a minimum document frequency of 10%. Using a min_df of 0.10%, both sets reached the highest F1-Score Macro, with values of 41.48% and 39.61%, respectively, and their means are statistically different. On the other hand, in Table 7, the ranges that exhibited the best performance were 2-5 and 2-6, once again in comparison to the reference, which consists of 2-5-grams with a minimum document frequency of 10%. Both sets, using a min_df of 0.10%, achieved an F1-Score Macro of 41.95% and 41.47%, respectively, and their means also differ significantly. In Table 8, it is noteworthy that the best performance was obtained with n-gram ranges of 1-4 and 1-5, as opposed to the reference, which is 1-2 grams with a minimum document frequency of 10%. These sets achieved an F1-Score Macro of 20.69% and 21.14%, respectively, and their means are statistically different as well. A notable observation across all three feature sets is that a minimum document frequency of 0.10% yielded the best results, while the feature set based on POS tags resulted in the lowest F1-Score values.
Macro F1-Score Average using 10-fold cross-validation with Word n-grams feature-sets (with n varying from 1 to 4 and min_df varying from 2 to 0.1%)
Macro F1-Score Average using 10-fold cross-validation with Word n-grams feature-sets (with n varying from 1 to 4 and min_df varying from 2 to 0.1%)
Macro F1-Score Average using 10-fold cross-validation with Characters n-grams feature-sets (with n varying from 1 to 5 and min_df varying since 2 to 0.1%)
Macro F1-Score Average using 10-fold cross-validation with Part Of Speech n-grams feature-sets (with n varying from 1 to 5 and min_df varying since 2 to 0.1%)
Finally, we used the feature sets that achieved the two highest performances for word n-grams (wng), character n-grams (cng), and POS tag n-grams (png), and merged them to generate a new feature set. Classification results using this feature combination are shown in Table 9. We can see that merging character n-grams (with n varying from 2 to 5) and unigram words produces the best results with F1 score 42.05%, and the means are statistically different using the combination of feature sets with the lowest value of F1-Score as reference. We conclude that character n-grams provide a better feature representation than POS tags and word n-grams individually.
Macro F1-Score Average using 10-fold cross-validation combining the best feature sets
The final experiments were conducted using deep learning pre-trained models based on the Spanish language, where we explored models primarily based on Bidirectional Encoder Representations from Transformers (BERT) and Generative Pretrained Transformer (GPT) used in the literature [4, 16].
The BERT multilingual case [7] was trained on the top of 104 languages with the largest Wikipedia content. This model is case sensitive 10. BERT Spanish [5] is a BERT model trained on a large Spanish corpus.
The models used a vocabulary of approximately 31k BPE subwords created using SentencePiece and were trained for two million steps. GPT2-spanish 4 is a language generation model pre-trained on 11.5GB of Spanish documents using a BPE tokenizer. Trained from scratch, the corpus includes 3.5GB of Wikipedia articles and 8GB of various books. The tokenizer was specifically designed for Spanish to capture semantic relations accurately.
Table 10 shows that BERT Spanish model consistently outperforms the bert-multilingual-cased model in all evaluated epochs. The GPT2-Spanish model generally obtains lower values compared to the BERT models. We believe this is because GPT-2 is a generative language model with an unidirectional architecture, which may not be as well-suited for the task of classifying communicative intentions as the bidirectional approach of BERT. Finally, we can see the difference using Deep Learning with respect to Logistic Regression contrasting the confusion matrices in Figs. 3 and 4. BERT Spanish got higher classification values for all communicative intentions except for the inform intention, where there was a 5% decrease in accuracy. However, the classification of the rest of the communicative intentions improved significantly. The clarify intention showed a 92% improvement; the stylish intention increased by 21%, the engage intention increased by 19%, the influence intentions increased by 16%, and emotional intention by 12%. All relevant data are within the paper and its Supporting Information files are available on 5
Macro F1-Score of Deep Learning Models
Macro F1-Score of Deep Learning Models

Bert Spanish Model Confusion Matrix.
We analyzed the most and least significant feature sets for each intention category. Subsequently, we conducted a comprehensive analysis of the classification results, encompassing both the machine learning and deep learning models.
Impact of n-grams in logistic regression model
To identify the most relevant features for each type of intent, we employed the coefficients of a logistic regression model. These coefficients reflect the weight of each feature (word or n-gram) on the probability of an instance being assigned to a specific class. To present these results more intuitively, we chose to use a word cloud as the basis for representing coefficient values. In this representation, the font size is determined by the coefficient value, allowing us to discern the most relevant features for each type of intent which are codified on different colors, as shown in Fig. 5.

N-gram Cloud of Top Positive Coefficients by Intention in Logistic Regression Model.
In Fig. 5 we can observe the following insights: The intention to influence can be conveyed through requests that include n-grams, such as “favor”, “por favor” (please), “espero” (I hope), “queremos” (we want), “escucha” (listen), and “no te” (don’t). Each n-gram is oriented towards the intention to influence, encourage an action, or elicit a reaction from the recipient in a specific context. The emotional intention is present in various forms, including positive expressions, such as “amo” (love), “ok” (okay), “buena” (good), “perfecta” (perfect), and laughter expressions like “jajajajajaja” (hahahaha) and “jajajajaja” (hahaha), as well as negative emotions, such as “triste” (sad). The essence of emotional intention is to convey what the addressee feels at a given moment, and this is reflected in the n-grams “yo” (I), “me” (me), “no me”, and “ay” (oh). The latter sequence expresses sympathy or concern for something mentioned. Additionally, we find n-grams that include symbols, such as “exclamacion pregunta” (exclamation question). This sequence combines exclamation and question marks, suggesting surprise or astonishment in a question. The intention to engage the interlocutor is manifested in the n-grams that include greetings, such as “saludo” (greeting) and “buenos dıas” (good morning). Furthermore, in relation to n-grams that incorporate symbols, the question mark “?” plays a fundamental role in maintaining interaction in the conversation, as evidenced in n-grams “?” (pregunta) and “no ?” (no pregunta’). In addition, expressions such as “!! cara” (exclamacion exclamacion cara), “!..” (exclamacion punto punto), “!!.” (exclamacion exclamacion punto), “..se” (punto punto se), and “..,” (punto punto coma) contribute to establishing an emotional connection in the conversation to keep the channel open. The clarify intention is reflected in the n-grams “nombre de’ and “suena’, which can be used to explain the correct pronunciation of proper names or specific words. Additionally, “escriben” and “como la” can be helpful in clarifying how certain words or phrases are written, especially when there are spelling variations or alternative forms. Finally, “como se” and “por que” are employed to address questions or concerns related to how certain actions are performed or the reasons behind certain phenomena. These n-grams were used to provide linguistic explanations and clarification. The stylish intentions are expressed artistically and creatively to create stylistic effects in communication. One of the most prominent n-grams is “es como”, which refers to a fundamental analogy in this intention. The intention to inform involves conveying information and establishing referential connections in conversation. Within this context, the n-grams “este es”, “por lo que”, and “que son” are expressions that can be used to affirm or confirm information in referential discourse. Similarly, “de el” and “en este” are employed to indicate the location or ownership of something in this referential context.
To analyze the classification results, we compared the performance of the principal models used in this research for every communicative intention, as shown in Fig. 6, in this upset plot and filtered the results from the intersection of the four classifiers to enhance visualization. At this intersection, which represents 51.9% of correctly classified instances (a total of 630), we observe that the categories with the highest number of correctly classified instances are intention to inform and emotional, followed by stylish and influence. It is worth noting that in the clarify intention, there was only one instance that was correctly classified by all classifiers; however, the engage intention was not present in this intersection.

Classification Results by Algorithm and Intention.
The incorrectly classified instances represented 10.5% of the total, with a total count of 127. When considering the count per category, we observe that the category with the highest number of incorrectly classified instances is the inform intention, followed by emotional, stylish, influence, and engage. We note that the stylish and engage categories are particularly challenging to classify, which aligns with the results of the confusion matrices shown in Figs. 3 and 4.
When evaluating Deep Learning algorithms in conjunction with individually obtained correct classifications by Machine Learning, we note that both BERT Spanish and BERT Multilingual made significant contributions to the influence, engage, and stylish categories. It is important to highlight that BERT Spanish shows substantial differences compared to the other classifiers, followed by BERT Multilingual. Additionally, BERT Multilingual successfully classified the highest number of engage instances, followed by BERT Spanish. A higher number of instances is observed at the intersection between Machine Learning classifiers and BERT Spanish, which contains 117 instances highlighted in pink, compared to the intersection with BERT Multilingual, which has 45 instances highlighted in purple.
The black bars on the left side of the graph represent the number of correct classifications for each model. These data reveal that BERT Spanish performed the best in terms of accuracy, followed by the Logistic Regression model, SVM and BERT Multilingual. Finally, when comparing the results by algorithm type, namely Machine Learning and Deep Learning, we noticed that Deep Learning algorithms achieve a higher number of correct classifications in the stylish category, both individually and in their intersections, highlighted in blue, compared to Machine Learning algorithms, highlighted in orange.
This article introduces a novel perspective in the study of intent on Twitter, which includes a prominent connection between the communicative functions introduced by Jakobson (1960) [13] and the interlocutor’s intent, as noted by some authors [2]. To address intention detection, we developed several automatic classification models that encompass: a) different pre-processing methods, b) various feature sets, and c) a variety of classifiers. Regarding pre-processing methods, we concluded that emoticons and symbols provide relevant information that enables discrimination between intent types in Machine Learning classifiers. Furthermore, after selecting the Logistic Regression algorithm because of its excellent performance in choosing the pre-processing method with the highest F1 Score, we observed that the combination of character n-grams (with n varying from 2 to 5) and word unigrams achieves the best performance. As for classifiers, although Logistic Regression has similar accuracy to the BERT Spanish model, the F1 Score is higher for the latter, indicating that BERT outperforms in classifying all intentions. Corpus labeling is a critical factor in generating automatic classification models. We have achieved substantial value in the Kappa Agreement, demonstrating the validity of our corpus as a robust tool for addressing this issue. Using a mapping of Jakobson’s language functions to the intents proposed by Arcand yielded better results than the five supertypes of intent presented by Radina [15], based on Habermas [9]. This suggests that the complexity of defining boundaries between intent types can influence results.
For future research, we propose considering classification as a multi-label problem, as intentions often have inevitable overlaps. We also suggest leveraging the polarity and priority tags from the RepLab2013 corpus and the intent label to address topics related to marketing strategies and customer relationship management. For example, if a tweet has a positive polarity and emotional intent regarding a specific product, considering user profiles (gender, age, etc.) could be crucial in managing customer relationships. Another potential application is monitoring reputation management, where negative sentiment may lead to a focus on the reputation dimension mentioned in a tweet. Additionally, it is relevant to explore intention detection concerning users who tend to disseminate fake news when they intend to inform or influence the recipient and measure its impact on reputation and risk management processes. Finally, we propose creating subcategories of intentions based on those identified in this study, which could further enrich the intention analysis on Twitter.
Footnotes
Acknowledgments
This paper has been supported by PAPIIT project TA101722 and CONAHCYT CF-2023-G-64. Also, the authors thank CONAHCYT for the computing resources provided through the Deep Learning Platform for Language Technologies of the INAOE Supercomputing Laboratory and the CONAHCYT scholarship program (CVU: 631563). G.B.E. is supported by a grant for the requalification of the Spanish university system from the Ministry of Universities of the Government of Spain, financed by the European Union, NextGeneration EU (Marıa Zambrano program, Universitat de Barcelona).
