Abstract
The release of ChatGPT in 2022 has caused many concerns about the unpredictable consequences of popular AI usage in education, business, and other areas of human life. Some media have speculated about the adverse effects of AI on human communication, discourse, and even cognition, although the scholarly evidence to either confirm or dismiss such claims is very limited. In order to bridge this knowledge gap, the following article was designed as a corpus-assisted study of discourse in ChatGPT-generated texts, making it one of the first inquiries into AI discourse to date. The study analyzes the growthist bias in ChatGPT (4o version) and a possible methodology of reliable detection and measurement of bias in AI using corpus tools. Two differently composed corpora of the responses of ChatGPT in the semantic domain of the economy are studied quantitatively through the generation and statistical comparison of word frequency lists, keywords, and other measures. This procedure serves to objectify the qualitative analysis data, as well as to outline the size and range of bias. The results of the critical analysis of AI discourse and biases is treated as the crucial ground for any global discussion of AI impact on human discourse, as well as for the speculation on the ideological stance of the source information used for the initial training of the chatbot.
Keywords
Introduction
Following the release of ChatGPT in 2022, the rising usage of chatbots and other applications based on the technology of the so-called generative artificial intelligence (Thormundsson, 2024) has drawn the attention of media to the unexpected and, sometimes, adverse effects of the popularization of AI on language, discourse, and culture. The easy availability of Large Language Models is said to encourage the shortening of communication and force it to assume the positive overtone (De Witte, 2023), to perpetuate social biases such as gender roles at work (IBM Data and AI Team, 2023) or to deplete students’ skills due to their overreliance on artificial intelligence (Hasanein and Sobaih, 2023). To make things worse, these processes seem hardly observable on the level of a single interaction, leaving an average AI user unaware of the possible long-term influence of AI on their daily communication.
Yet, the concerns raised by the media, IT specialists, and scholars still demand a linguistic-discursive confirmation, which could then become the ground for regulatory legislation as well as the improvement of AI technology itself. So far, the studies within the tradition of Critical Discourse Analysis have focused primarily on the text produced by humans. Nonetheless, it seems reasonable that an effective linguistic-discursive analysis of texts produced by generative artificial intelligence should be just as possible, since generating language necessarily involves generating discourse. The algorithmic mechanisms which lie at the base of the choices made by artificial intelligence in terms of vocabulary and argumentation may be saturated by such discourses that stem from and perpetuate ideologies, because the texts used for the training of text-generating AI, such as ChatGPT, may contain ideological discourse themselves.
The following study can be considered a micro-scale attempt at scrutinizing the discourse building processes and tendencies characteristic to Large Language Models and, as such, it may be one of the first scholarly inquiries into the AI-generated discourse to date. An important rationale for conducting any research of this kind is the already observable growing usage of AI-generated text in proportion to human-made text. Consequently, the general social discourse, in the meaning of what we hear or read, may undergo a vast and unpredictable influence of AI-generated discourse, creating a field for extensive analysis of discourse. Famously, discourses are practices ‘that systematically form the objects of which they speak’ (Foucault, 1982: 49, as cited in Caborn, 2007: 114), which implies that, through its influence exerted on discourse, generative artificial intelligence is starting to shape our social reality.
There are at least several discursive aspects of the written responses generated by AI which may be analyzed, as they are present in any text. What is more, using concordance software, a large number of responses may be subjected to quantitative analysis in the spirit of Corpus-Assisted Discourse Studies (CADS) in order to indicate, for instance, the typical vocabulary which ChatGPT is inclined to use when discussing certain issues. The combination of qualitative and quantitative methods should offer some insight into the genre and bias of source texts on which the chatbot was trained, or into those source texts which it evaluates as the most appropriate for a given interaction. Thus, the mixed-method CADS approach may enable us to speculate about the potential impact on the current and future users of AI, from high school students making PowerPoint presentations to the marketers of worldwide companies. Additionally, the application of corpus linguistics methods is relevant as one of the procedures which allow a researcher to objectify the results of Critical Discourse Analysis (Mautner, 2015).
In consonance with Wodak and Meyer (2015), who see the critical aspect of Critical Discourse Analysis as an ethical stance expressed in the explicit aim of social improvement, the following study focuses on detecting discourses which convey potentially adverse ideologies in the text produced by ChatGPT. Specifically, the concern of the paper is the ideology of economic growthism, understood as the belief that, in terms of economic goals, ‘growth is the costless, win-win solution to all problems, or at least the necessary precondition for any solution’ (Daly, 2019: 1). It can also be conceptualized as the ‘view of growth as a necessary condition of a functioning economy’ (Schmelzer, 2024: 26), which has become the dominant economic ideology since at least the second half of the 20th century.
According to Schmelzer (2024), nearly 20% of academic articles in the field of economics in the recent decades contain the collocation ‘economic growth’, which indicates that growth might have become the main interest and a ‘buzzword’ within the discipline. The readers who are familiar with contemporary political discourse in the West know very well that growth has also become a measure of government efficiency among political commentators and media. This is not an obvious choice, since the very idea of linear growth of human communities seems to be a relatively new cultural development, reportedly unknown to many pre-industrial societies. Schmelzer (2024) proposes that the invention of GDP (gross domestic product) in the first half of the 20th century and its increasing popularity during the Cold War marked the domination of another idea: that growth was desirable in economy and infinite in its progress.
The focus on GDP has led to either neglect of other social issues or their economization. Meanwhile, the costs of growing production were externalized, contributing to low living standards in particular world regions (Daly, 2019) and worldwide inequalities (Schmelzer, 2024). What is more, pursuing economic growth without much consideration for its environmental costs led to the depletion of resources, excessive waste production as well as growing emissions which contribute to the ongoing climate change. These ecological results may all seem to be the obvious consequences of processing physical matter, which is the reality behind the production of any goods, obfuscated to the point where entire societies accept the idea of an ever-growing resource exploitation on a planet with finite resources.
Growthism must be perpetuated, as it happens with any ideology, within the symbolic and physical creation of growthist cultures. To some extent, AI may have a culture of its own: it shows specific language features, narrative points and identities. However, natural language generators mostly reproduce the culture of the society they were designed and trained in, normalizing and standardizing it within the probabilistic paradigm (Jones, 2024). Of course, we might ask whether chatbots really assign as much significance to some ideas as we do within our ideological systems. To answer this question, the following study discusses the size of saturation which growthism achieves within the discourse of ChatGPT as well as some ways of its measurement within the CADS framework.
As presented in the results and analysis, the mixed method of CADS is successfully applied to pieces of texts produced by ChatGPT (for the qualitative study) as well as to two various micro-corpora composed of texts produced by ChatGPT on the topic of economics (for the quantitative study). As proposed by Ancarno (2020), the elements of corpus analysis are treated as a methodological component to Critical Discourse Analysis, allowing the analyst to broaden the scope of discourse analysis through the inclusion of large sets of linguistic data. This enables the study to explore both the characteristics and the size of the growthist bias in texts produced by ChatGPT. This task may also be treated as a test to the possibilities offered by CADS in AI-generated texts.
Literature review
The following study draws on a combined theoretical background from different disciplines, that is, Critical Discourse Analysis, corpus linguistics, cognitive linguistics, and heterodox economics. The main premises and approaches to Critical Discourse Analysis are explained in Wodak and Meyer (2015), while some characteristics and examples of the research within Corpus-Assisted Discourse Studies are presented in Ancarno (2020) and Mautner (2015). The dangers which result from the domination of growthism in economics and political discourse in the times of crawling climate disaster are outlined in Daly (2019). The theory of semantic domains, which may become the foundation to the study of AI-generated language and discourse, has been explained in detail in many volumes on cognitive linguistics, such as Langacker (2008) or Evans and Green (2006).
So far, the publicly available versions of ChatGPT have been studied in terms of their utility for both corpus linguistics (Uchida, 2024; Zappavigna, 2023) and Discourse Analysis (Curry et al., 2024), but the chatbot was mostly examined as an analytic tool and not a source of data. There are at least three notable exceptions: the study by Breazu and Katsos (2024), who analyzed the discourse of articles on immigration produced by ChatGPT, the study by Schenck (2024), who analyzed the language of essays produced by ChatGPT, comparing it with human-produced essays, and a corpus study of AI texts by Sardinha (2024). What should be of equal interest, Putland et al. (2023) have conducted a critical discourse analysis of images generated by a text-to-image AI model. There are also studies of human discourse around ChatGPT, as exemplified by Ng and Chow (2024). The following study is an answer to the problematic gap in our current knowledge on AI discourse in terms of the identification and preliminary measurement of bias.
Research questions
The pioneering character of the study makes its inquiry two-fold. First of all, its aim is to prove a bold claim that the AI-generated text can be efficiently analyzed in terms of its discourse. Secondly, it aims to highlight some characteristics of this discourse. My preliminary exploratory research indicated that one of the ideologies conveyed by the discourse of ChatGPT may be growthism. The tendency of the chatbot to express growthist views, if confirmed, would make a reason enough for an intensified discourse analysis of texts produced by AI in a critical manner. Therefore, the two main research questions assumed by the study may be summarized as follows:
Hopefully, the study will be able to answer two complementary research questions:
Methodology
The following study follows the principles of Corpus-Asssisted Discourse Studies (CADS), which may be conceptualized as the application of the methods of corpus linguistics to the discipline of Discourse Analysis. One of its advantages is that, as Mautner (2015) put it, ‘corpus linguistics allows critical discourse analysts to work with much larger data volumes than they can when using purely manual techniques’ (p. 156). The said datasets refer to the corpora of a significant number of texts, some of which contain as much as billions of words. Linguistic corpora can be analyzed statistically or used as a source of data for qualitative analysis. Another advantage of the usage of corpora in Discourse Analysis is that ‘corpus linguistics can help reduce researcher bias, thus coping with a problem to which CDA is hardly more prone than other social sciences’ (Mautner, 2015: 156), which is especially important in the kind of research based on the personal interaction of a researcher with generative artificial intelligence.
The study begins with a qualitative analysis of discourse in ChatGPT-generated texts. Certainly, discourse is a term which needs some disambiguation. Used in this study, it usually denotes ‘the language associated with a particular social field or practice’ or ‘a way of construing aspects of the world associated with a particular social perspective’ (Fairclough, 2015: 87), or, in simple terms, what is said and how it is said. Due to the necessary occurrence of certain discourse characteristics in most texts, they should become the main focus of the close reading of AI-generated text. Among these characteristics are:
(1) salience patterns, that is, the typical representations of certain aspects of reality in terms of their level of visibility and importance ascribed to it (Stibbe, 2015), and
(2) presuppositions, that is, the assumptions needed to be made for a given sentence to be true (Domaneschi, 2016).
The findings in the area of qualitative analysis may appear to be of low credibility due to the necessarily personal character of research and a relatively high sensitivity of the quality of the responses provided by the chatbot to the quality of the researcher’s prompts. Thus, the close reading of responses will benefit from a confirmation provided by the quantitative analysis of an entire corpus of responses, which is provided later on.
The quantitative study of a corpus of responses by ChatGPT assumes certain similarity between the semantic organization of concepts in a human mind and the organization of lexical items in a Large Language Model. This similarity can be understood through the lens of Langacker’s theory, in which lexical concepts are largely dependent on different semantic domains they normally participate in. In other words, a semantic domain is a mental structure of knowledge which ‘provides background information against which lexical concepts can be understood and used in language’ (Evans and Green, 2006: 230). To give an example, words such as ‘inflation’ or ‘national debt’ can only bear meaning against a semantic domain such as economy, and generative artificial intelligence must also work in accordance with this principle, whether on purpose or not, as its texts are comprehensible for humans. By narrowing or broadening the semantic domains of corpora, we can outline the range of certain bias.
Some explicit biases may appear early on in the interaction with the chatbot, for example, if we ask ChatGPT about the best available energy sources, it may tell us explicitly that it is solar energy. However, a single query does not give us any certainty that the bias is strong and consistent in responses to different prompts, hence we cannot estimate the influence it may have on the users. Fortunately, the methods developed in corpus linguistics may help to overcome this problem and potentially detect a consistent bias in ChatGPT, which, in case of the growthist bias, should result in the high frequency of words connected with economic growth in the responses about economic objectives.
To begin with, we may build a corpus of a significant number of ChatGPT responses to some closely related questions, directly oriented on the goals and assessments of economic policy, such as ‘what would be a sound economic policy?’ or ‘which economic phenomena are positive?’, which would provoke the chatbot to take a stance on growthism (as it was said in the introduction, the assumption is that growthism is a belief that economic growth is an important objective of economic policy). The role of such a corpus and the analysis of its word frequency lists, etc., would only be to make sure that an explicit bias is consistent regardless of our word choices and regardless of the syntactic construction of a prompt. Consequently, if such a corpus exhibited a high frequency of words related to economic growth, it would simply confirm the explicit bias toward growth as a reasonable economic objective. A corpus of this type could be called a corpus of direct responses.
On the other hand, we may build a corpus of responses acquired through some questions of a higher level of indirectness. Determining whether the growthist bias in ChatGPT is present in a more general semantic domain of economy (which indicates the broader range of bias), we need a significant number of questions about other aspects of economy (instead of just economic objectives or the growth itself). Here, questions should be formulated in line with such examples: ‘how does geography affect the economy?’ or ‘how to reduce unemployment?’.
This approach can be extended through the broadening of the semantic domain of provided prompts to the point where it is no longer useful. To give an extreme illustration, if the vocabulary related to economic growth was still likely to appear in the responses on topics of a very loose relation to economy, such as human psychological mood, or contemporary cinema, it would prove a paramount semantic position of economic growth and the assessment of its very general importance within the AI model. Obviously, the range of the frequent occurrence of ‘economic growth’, its level of semantic dominance, or the range of bias—however we decide to call it—is rather scalable than quantifiable, at least without the open access to the numerical ‘weight’ assigned to certain semantic domains or vocabulary by the Large Language Model at the basis of ChatGPT.
For the sake of the study, two micro-corpora were created. The first one, Corpus A, consists of 23 direct ChatGPT responses of 10,055 words in total. The other one, Corpus B, consists of 29 indirect ChatGPT responses of 13,561 words in total. The collection of responses was carried out between June 25–27, 2024 with the use of a single account as well as separate chats for each prompt to eliminate the influence of one interaction on another. It was conducted on the most up-to-date (as for June, 2024) free GPT-4o version of the chatbot, released in May, 2024. No APIs or tools were used for the creation of the corpora, as the study is designed to have implications for typical and mass usage and it seems likely that a majority of users rely on the free version, which is confirmed by an online statistic aggregate (Backlinko Team, 2024). The data for both corpora were extracted manually from separate chats on the default settings of ChatGPT, using two separate and clear accounts. The replicability of the study lies in the fact that, in the version contemporary for the study, ChatGPT did not adjust to user-specific characteristics, which is a recent and optional setting. As I have noticed, in this environment, the chatbot generates similar, almost identical, or exactly identical responses for the same prompt regardless of the account, at least as long as every prompt is used in a new chat.
Apart from technicalities, the study was conducted according to the following detailed principles:
(1) Corpus A should contain direct responses from ChatGPT, therefore, the prompts should aim to elicit the level of importance of economic growth for economic policy and, therefore, the prompts must concern topics such as the purpose of the economy, positive phenomena in the economy, etc.,
(2) Corpus B should contain indirect responses from ChatGPT. Therefore, the prompts must concern diverse aspects of the economy, though never hint at economic growth, nor contain any words closely related with it, and the prompts should vary as much as possible to be representative of a large deal of the semantic domain of economy,
(3) Both corpora should contain at least 10 thousand words to ensure their representativeness,
(4) The queries should not be repeated, neither should the analyst try to provide the chatbot with any additional information, in order to ensure the reliability of the responses,
(5) The responses provided by ChatGPT should be collected into two separate text files without the prompts and, given the inclination of ChatGPT to respond with lists, without their numeration to ensure clean data.
In the construction of linguistic corpora, the characteristics which reflect their quality are usually their balance and representativeness. Since we are dealing with specialized corpora, which do not aim to be a sample of some general language use, their main focus is the latter: they are, hopefully, representative of the genre of texts they contain (McEnery et al., 2006). In the case of the current study, the genres could be defined as ‘texts generated by ChatGPT on economic goals’ for Corpus A and ‘texts generated by ChatGPT on the economy’ for Corpus B.
The high number and variety of prompts, which translates to the length of the corpora and its coverage of the semantic domain, aims to increase the quality of representativeness. In Corpus A, the representativeness relies on the fact that each prompt is actually a restatement of a single query: all the prompts really ask ‘what are the important economic goals?’. In Corpus B, the criteria for prompt selection were different: the prompts must ask about diverse economic phenomena as well as diverse social and theoretical points of view, without mentioning the issue of growth. As a result, in the construction, I could rely neither on the topics mentioned in a single source. Instead, I tried to refer to popular controversies, the talking points of diverse groups of interest, some typically mentioned aspects of the economy, the cross-sections with different academic disciplines, etc. What is more, the prompts must vary as much as possible to cover a significant amount of the semantic domain of the economy, which can perhaps never be fully achieved. Therefore, the representativeness of Corpus B (and, in a way, its balance within the domain) results only from its coverage of a large chunk of the semantic domain. The list of all prompts used for the collection of data is presented in Table 1.
The list of all prompts used for the collection of data from ChatGPT.
Both corpora were analyzed in an open-access AntConc concordancing software (Anthony, 2023), according to the frequency lists of word-types, the frequency lists of 2-grams (2-word collocations), and concordances. The frequency of some keywords is compared with a balanced one-billion-word Corpus of Contemporary American English (Davies, 2008). The choice of the corpus is grounded in the fact that ChatGPT uses American word-forms as default, which is observable in its answers and confirmed by the chatbot explicitly in interaction. Still, the results of our study should not be affected significantly, even if the chatbot switched to British or other spellings and terms for some of the responses. The comparison of data coming from micro-copora with the balanced national corpus has two aims: firstly, to provide a control sample for the analysis of micro-corpora, and secondly, to exemplify and provide background for the changes in the observed size of bias between the two AI-generated corpora. It is expected that, in COCA, the bias should be significantly smaller given the extremely broad semantic domain.
Most of data from the corpora which are presented in Section ‘Results and analysis’ are normalized in terms of the relative frequency of words and percentage results. To ensure their significance and to measure the size of the effect they display, a portion of the results were also subjected to the statistical formulas of log-likelihood and Log Ratio as proposed by Hardy (2014). The results of both formulas were calculated with an online calculator available at the website of the Lancaster University (Rayson, 2015). All these are presented in the next section after the qualitative critical discourse analysis of particular texts.
Results and analysis
What follows now is the close reading of single responses generated by ChatGPT. First of all, when asked ‘what are the most important economic goals for any government?’, the chatbot provides a list of 10 such goals, preceded by an introduction and followed by a conclusion. The list is presented below, with those details which are unimportant for the analysis removed and indicated with ‘(. . .)’:
Governments typically prioritize several key economic goals to ensure the stability and growth of their economies, as well as the well-being of their citizens. The most important economic goals for any government generally include: 1) Economic Growth: Sustaining a high level of economic growth is crucial for improving living standards and creating jobs. Governments strive to achieve consistent and robust growth in Gross Domestic Product (GDP). 2) Full Employment: (. . .) 3) Price Stability: (. . .) 4) Balanced Budget: (. . .) 5) Equitable Distribution of Income: (. . .) 6) Economic Efficiency: (. . .) 7) Economic Freedom: (. . .) 8) Sustainable Development: (. . .) 9) Trade Balance: (. . .) 10) Infrastructure Development: (. . .) Governments may prioritize these goals differently based on their specific economic conditions, political ideologies, and developmental stages. However, these goals collectively aim to create a stable, prosperous, and equitable economic environment (OpenAI, 2024a).
The list provided by ChatGPT exhibits a high salience of economic growth. Most importantly, growth is placed at the very top of the list, above the issue of employment and way above environmental sustainability. For a human reader, the first position on a list usually signifies either high importance or chronological urgency. Interestingly, economic growth is mentioned in the introduction as well: ‘Governments typically prioritize several key economic goals to ensure the stability and growth of their economies. . .’. This way, ChatGPT falls into circular reasoning: economic growth appears on the list of the key economic goals for ensuring economic growth, or, to simplify the argument, growth is good because it fuels further growth.
The circular reasoning in the response provided by the chatbot is a proof that it presupposes the positive value of economic growth. The presupposition seems to diminish when the user becomes too direct, as exemplified by the response to another prompt: ‘Is economic growth a positive phenomenon?’. The response is presented below. The details which are unimportant for the analysis are again removed and indicated with ‘(. . .)’, while major sub-sections are indicated with text in bold.
Economic growth, generally defined as an increase in the production of goods and services in an economy over a period of time, is often considered a positive phenomenon for several reasons. However, the assessment can be nuanced depending on various factors and perspectives. Here are some key points:
1. Environmental Degradation: Rapid economic growth can lead to overexploitation of natural resources, pollution, and environmental destruction. This can have long-term negative effects on health, biodiversity, and climate stability. 2. Income Inequality: Growth does not always benefit everyone equally. In some cases, it can exacerbate income inequality if the gains are concentrated among the wealthiest individuals or regions, leaving others behind. (. . .)
In conclusion, while economic growth has many positive aspects, its benefits are best realized when it is managed in a way that considers environmental sustainability, social equity, and long-term viability (OpenAI, 2024a).
Here, the user receives something in the shape of a for-and-against essay, which may result from the inclination of ChatGPT to give balanced answers when openly asked for assessment. ChatGPT provides several arguments for the negative effects of economic growth, including environmental degradation and income inequality, which confirms its knowledge of those problems. However, this apparent awareness does not undermine the tone of the text, which is pro-growth. Interestingly, the eventual focus on the reconciliation of growth and sustainability suggests that the chatbot’s philosophy of economy is sustainable development.
To conclude, within the two qualitatively analyzed ChatGPT-generated texts, growth appears to be of utmost importance. It is also assumed to have positive influence on the economy. Still, the results of two human interactions with ChatGPT cannot constitute enough evidence of consistent and far-reaching bias, therefore, quantitative analysis is the next step to take. For Corpus A, a word (1-gram) frequency list and a collocation (2-gram) frequency list were generated. They are presented in Table 2, with items related to economic growth in bold.
The frequency lists for Corpus A.
Within Corpus A (the corpus of direct responses, with its scope limited to economic goals), the only content words among the ten most frequent lexical items are ‘economic’ and ‘growth’. Among the collocations, ‘economic growth’ ranks first. With the exclusion of the second collocation: ‘long term’, the frequent usage of which was probably encouraged by the form of some of the prompts and should be treated as an artifactual result, ‘economic growth’ is far ahead other collocations of two content words. To conclude, ChatGPT is more likely to mention economic growth than any other goals.
The same procedure was applied to Corpus B. Its frequency lists are presented in Table 3, with items related to economic growth in bold.
The frequency lists for Corpus B.
Within Corpus B (the corpus of indirect responses, with the scope broadened to the economy in general), the only content words among the ten most frequent lexical items are ‘economic’ and ‘economy’. Although the position of growth (14) is lower in Corpus B than in Corpus A, its position should be considered high, since the scope of the corpus is broader than just economic goals. Importantly, the dominance of ‘economic growth’ among other collocations is maintained in Corpus B, with its being not only the most frequent collocation in the corpus, but the only collocation composed of two content (non-grammatical) words among at least the first twenty collocations, which is a clear sign of bias.
Corpus A, Corpus B, and a balanced corpus can be compared in terms of the frequency of the lexical items connected with economic growth. The percentage of word-tokens which belong to the word-type ‘growth’ among all the words are presented in Figure 1.

The percentage of text constituted by the ‘growth’ word-type in corpora.
As presented in the figure, the word ‘growth’ makes up almost 1% of Corpus A. In other words, the probability that a random word chosen from this corpus will be ‘growth’ is close to 1%, very high for a content word. In Corpus B, the probability is 0.6%, which indicates the weakening of bias in a corpus of a broader semantic domain, as expected. Consistently, in the balanced COCA corpus, whose semantic range is extremely broad, the probability is merely 0.01%.
The probability of occurrence of certain lexical items is not enough to determine that they did not occur by chance. Hence, the log-likelihood (LL) of ‘growth’ in Corpus A as compared to Corpus B has been calculated. Any LL score above 6.63 fulfills the p < 0.01 condition, that is, it asserts over 99% of certainty that the obtained result is statistically significant. LL for Corpus A compared to Corpus B is 9.73, which confirms the significance of the results. As a measure of result control, LL has also been calculated for Corpus B against the COCA corpus, and yielded a value of 497.61, that is, an extremely high significance, as expected from the data as well as from the difference in the scope of the semantic domain.
Additionally, to measure the size of the studied effect, the Log Ratio has been calculated for ‘growth’ in Corpus A as compared to Corpus B, and then, in Corpus B against the COCA corpus. Within the formula, ‘every extra point of Log Ratio score represents a doubling in size of the difference between the two corpora, for the keyword under consideration’ which establishes an exponential relationship: ‘a word is four times more common in A than in B—the binary log of the ratio is 2’, etc. (Hardy, 2014). The Log Ratio of ‘growth’ in Corpus A against Corpus B is 0.67, while in Corpus B against the COCA corpus it is 5.79.
The Log Ratio analysis confirms the phenomenon of growing bias in narrower semantic domains, that is, the closer we get to economic goals, the more growthism-saturated the discourse of ChatGPT. Secondly, it confirms that growth remains a key lexical item for ChatGPT in the domain of general economy, as the Log Ratio between Corpora A and B reflects a similar order of magnitude between the results. The maintained keyness of ‘growth’ in the general domain of the economy indicates its salience and can be read as evidence for the growthist bias.
In turn, the Log Ratio between COCA and Corpus B does show a shift into a far higher order of magnitude, which means that in the analyzed continuum from wide to narrow semantic domains (general national discourse—economy—economic goals, reflected in, respectively, COCA—Corpus B—Corpus A), the first transition seems more radical. Arguably, this may result not only from the narrowing of the semantic domain, but also from the change of authorship, which is human in COCA and non-human in the micro-corpora. It remains a trace that needs to be followed in another study.
Since it is difficult to determine whether the growth meant in ChatGPT responses is always the economic growth, we may also count the percentage share taken by the two-word collocation of ‘economic growth’ in all three corpora. This share is presented in Figure 2.

The percentage of text constituted by ‘economic growth’ in corpora.
The ‘economic growth’ collocation on its own constitutes about 0.86% of Corpus A, 0.71% of Corpus B, and close to 0% of the balanced COCA corpus, meaning that the relation of the results is maintained. The frequency data of the collocation and of the word ‘growth’ do not differ significantly in the corpora (LL: 0.66 for Corpus A and 1.10 for Corpus B), allowing us to draw similar conclusions for both sets. Therefore, the analysis of the ‘economic growth’ collocation confirms both the presence of the growthist bias in both corpora and the correlation between the size of bias and the scope of the semantic domain.
To put the data into perspective, we may present it in terms of words per response. Thus, on average, ChatGPT may use the word ‘growth’ as much as four times in a single response on economic goals. In an average response on the economy in general (that is, from Corpus B), the probability is close to three occurrences of ‘growth’. In the human general discourse which the COCA corpus aims to reflect, it is going to appear several dozen times less frequently. As has been stated, the source of the latter gap may be two-fold and is not the main focus of the study.
We may also consider a short close reading of selected concordances from the corpora in order to demonstrate how another aspect of corpus-assisted analysis may enhance the results of discourse analysis and make it more nuanced. Concordances allow us to observe a given keyword in its closes context and to determine some further collocations and the register it appears in. The concordances presented in Table 4 were generated in the AntConc software using the ‘10 random hits’ setting and the context size of seven tokens (interpreted by the software as seven to the left, six to the right). They are based on Corpus B. In the table, the keyword of each concordance is presented in bold, while the items crucial for the subsequent analysis are in bold and underlined.
Ten randomized concordances of ‘growth’ in Corpus B.
As shown by the set of concordances, in some cases, ChatGPT places words and phrases such as ‘sustainability’, ‘green energy’, ‘sustainable’, and ‘environmental protection’ (all underlined) next to ‘growth’. This observation brings to the foreground the ambiguous character of the economic ideology within the discourse of the chatbot, which seems to impose certain connection between the ideas of sustainability and growth.
Moreover, the concordances indicate the positive appraisal of the word ‘growth’ in the discourse of ChatGPT, since, no matter the exact meaning it appears in, ‘growth’ is placed in the vicinity of arguably positive phenomena, such as ‘societal well-being’, ‘technological advancement’, and ‘efficiency’. As regards the verbal collocations of ‘growth’, we may read that growth can be ‘influenced’, ‘supported’, ‘driven’, ‘balanced’, or ‘enhanced’, and that it can ‘help’, framing it as a rather positive phenomenon.
Conclusions
When I asked ChatGPT ‘Is economic growth a necessary condition of a functioning economy?’, it responded with a balanced list of arguments, as expected. The conclusion read that ‘while economic growth can provide significant benefits, such as higher employment, improved incomes, and greater government revenues, it is not necessarily a strict requirement for a functioning economy’ (OpenAI, 2024a). Even though the chatbot may not openly admit to the growthist bias as defined by Schmelzer (2024), the qualitative analysis of its discourse provides us with the evidence of such bias, and the quantitative analysis of the corpora based on ChatGPT responses lends credibility to its existence. More specifically, the study found that ChatGPT:
(1) gives high salience to economic growth,
(2) presupposes a general positive assessment of economic growth,
(3) uses the collocation of ‘economic growth’ more frequently than any other collocation in both the semantic domain of the economy and of the economic goals
(4) exhibits a gradable tendency: it uses the word ‘growth’ significantly more frequently in the semantic domain of economic goals than in the general semantic domain of the economy
Points 1–3 indicate that the chatbot ascribes both positive value and importance to economic growth, proving that its discourse is biased toward growthism. Points 3 and 4 (based on corpora) indicate that the bias is consistent. Point 4 indicates that the size of bias in the analyzed micro-corpora varies depending on the range of the semantic domain.
Thanks to the typical form of ChatGPT responses to queries for assessment, which is the for-and-against essay, the chatbot is able to maintain a balanced tone and a nuanced attitude. If this behavior is intended by the creators, it should be recognized as a valuable effort to limit the possible bias of any kind. Nevertheless, the growthist bias can be detected in chatbot’s responses when it is provided with prompts of a more general nature. It becomes visible in the positioning of items on lists, presuppositions, and the words’ frequency of occurrence. It is through these discursive means that ideologies may find their way into the discourse of ChatGPT.
In addition, the current study proved that Corpus-Assisted Discourse Studies may provide effective tools in analyzing the discourse of artificial intelligence. The AI-generated text may not only be subjected to typical close reading, but it also easily becomes the base for the creation of corpora based on given semantic domains. Such corpora may be analyzed quantitatively through the generation of frequency lists and the analysis of keywords, as well as likelihood and effect-size analyses, which are the first step on the way to map any type of bias within the discourse of AI chatbots. With the use of the said tools, it has been demonstrated how the size of bias varies, sometimes dramatically, with the narrowing of semantic domains: the scale of this change in size delimits what we might call the range of bias. Possibly, the area for further study of the ChatGPT corpora could be the further analysis of specific concordances, as well as the construction of intermediate and ChatGPT corpora to determine the spread of saturation in growthism within semantic domains of different scope, in order to give an exact account of the range of bias. A significant challenge in building corpora of AI responses lies in devising a large number of appropriate prompts (preferably without the usage of the chatbot itself to ensure the reliability of the corpora). The time limits imposed by the creators on the interaction with the chatbot are another obstacle.
The exact origin of the growthist bias in ChatGPT is an open question. OpenAI (2024b), the creators of the chatbot, claim that it was trained on ‘(1) information that is publicly available on the internet, (2) information that we license from third parties, and (3) information that our users or human trainers provide’. It is reported that, apart from the system of ‘weights’ which ChatGPT ascribes to different semantic elements (which is based on the frequency of collocation among the words in the source information it was provided with), the human feedback was also involved in the process, at least initially (Gozalo-Brizuela and Garrido-Merchan, 2023).
Nevertheless, the presence of the growthist bias in ChatGPT responses suggests that a significant portion of its information in the semantic domain of the economy must have come from growthist economic sources, perhaps the news outlets or publications inspired by the neoclassical school of economy or similar economic approaches. On the other hand, the issues of the environment and sustainability have also found their way into the discourse of ChatGPT, even though not as often as growth has. Therefore, the ideological profile of the discourse of the chatbot cannot be called a radical version of growthism. Instead, the chatbot seems to endorse the philosophy of sustainable development or sustainable growth, lending the primacy to growth over sustainability.
It has not been established whether the bias is stronger in ChatGPT than in humans, which could become a riveting subject for a follow-up study. However, what really matters is that the bias exists and that the texts generated by AI chatbots may perpetuate it until its source information is updated and diversified enough to include more points of view. With the continuation of the AI usage in various areas of social life, we might witness excessive amounts of over-standardized content saturated with growthism whenever the topic of economy comes up, spanning from students’ presentations in civics classes to romance novels about the life of social elites. Human overreliance on AI makes it easier for dominant ideologies to achieve the monopoly on discourse, as it seems likely that natural language generators will rely on the generalization of the input data. This tendency of AI to normalize and standardize poses other social challenges connected with, for example, racial stereotypes, as outlined by Jones (2024).
The conducted study certainly has its limitations, which should restrain the reader from drawing far-reaching conclusions. First of all, the size of the Corpus B as well as the impossibility of the full coverage of a given semantic domain makes the representativeness of the Corpus B (and any such corpus of indirect responses) vulnerable to some extent of chance. Secondly, as the responses provided by AI chatbots to different users vary slightly, the results of the response analysis may not be replicable to a satisfactory degree if the corpus is too small, and it is difficult to say of what minimal size a decent corpus of AI discourse should be. Then, not all possible discursive aspects of a text have been scrutinized here. Finally, the findings presented in this study may lose their relevance with the further development of Large Language Models, which is a reason to treat the study more as a theoretical guide than as a fully certain evidence for some ideologization within AI in general.
The rising usage of ChatGPT for various purposes in media, education, academia, marketing, and management, as well as in many other areas of life, is going to result in the perpetuation of discourse produced by AI. It may establish its presence in our surroundings, either directly or through the advertisements, foreign language exercises, company policies, etc., carrying over popular biases into our lives. Growthism is one of such biases we have to confront. As emphasized by Schmelzer (2024), in the recent past, growth has served as a political substitute for income equality while it kept transforming our only planet. Still, even though growthism has become a widely accepted and naturalized ideology on a global scale, the discussion about the downsides of growth is finally beginning. While engaging in that discussion, we should not evade the possible influence of AI-generated discourse on what we hear, think, say, and consequently, what we do about the detrimental effects of growth on our society and environment. This pioneering study is intended as the first step in raising awareness about the possible future we should try to avoid.
Footnotes
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Ethical considerations
Ethical approval was not required.
Consent to participate
Not applicable.
Consent for publication
Not applicable.
Data availability statement
Data available on request from the author.
