Abstract
This was an invited paper meant to accompany a paper by Bill Cope and Mary Kalantzis.
C&K
Bill Cope and Mary Kalantzis (hereafter C&K) offer a distinctive perspective on “Generative AI.” It is important to realize that Generative AIs like ChatGPT-4 (hereafter just “Chat”) have their own views on C&K’s claims. This is surely a unique situation in history. When we critique the design of a car, book, machine, or movie, we do not expect them to be able to reply.
So, before we introduce our own perspective on Generative AI, let us, momentarily, turn this paper over to Chat. We have truncated Chat’s answers due to space limitations. C&K may be happy to find out that Chat, by and large, agrees with them. When we get the paper back, we will disagree with both C&K and Chat on some points.
GPT-4 Turbo
Do you use grammar when you generate language?
I don’t use a traditional grammar book or set of explicit rules like a human might when learning a language. Instead, I am based on a machine learning architecture known as a transformer, which is trained on a vast corpus of text data. This training process involves learning patterns of language, including grammar, syntax, semantics, and common usage, directly from the dataset… My “understanding” of grammar is implicit, derived from the patterns in the data on which I was trained.
How do you handle semantics—the meanings of words?
Here’s a simplified explanation of how I deal with semantics: 1. 2. 3. 4. 5.
Is your method “anti-grammar” (as C&K claim)?
The claim, overall, reflects a valid critique of the limitations of statistical language models from a semantic perspective. It highlights the fact that while models like mine can produce and interpret language in ways that often seem meaningful to humans, the underlying process is devoid of true semantic understanding or intentionality.
Since humans have massive knowledge of context when they speak—and are not consciously processing language, for the most part, when they speak, how do we know they are not choosing the most probable word, given the context, when they speak?
Indeed, there is a theory in psycholinguistics known as the “usage-based model” that suggests that language learning and production are heavily influenced by frequency and probability based on experience. According to this model, humans, particularly in fluent conversation, often choose words and construct sentences based on patterns they have encountered before, which could be seen as selecting the most probable linguistic elements given the context.
Do you think systemic Functional Grammar can be applied to other modes than language and even to ChatGPT as a multimodal form of AI?
Systemic Functional Grammar (SFG) is fundamentally a linguistic theory, designed to analyze language in various contexts. However, the principles of SFG, particularly the focus on the communicative function of language, have been extended to multimodal discourse analysis. It’s important to note, however, that SFG was not designed to analyze the internal mechanisms of AI or any computational processes. It is more about the output and how that output functions within a communicative context.
Our perspective: The brain as a prediction engine
We will argue that Chat operates more similarly to humans than C&K and many others believe. A massive amount of human thought and action is carried out unconsciously (Barnett, 2020; Gazzaniga, 2011; McGilchrist, 2021). These unconscious processes include a good deal of perception, memory, learning, decision-making, and language production. Brain imaging techniques have revealed that decisions can be predicted based on neural activity before individuals are consciously aware of making a choice (Koenig-Robert and Pearson, 2019). In many cases, the idea we freely and consciously make a choice is an “after-thought.”
As in the case of Chat, the human unconscious brain can do its job well only if it has very “big data.” The brain is constructed, for the most part, from experience (Barnett, 2020; Eagleman, 2020). Each experience we have in life connects or disconnects neurons in our brain. At any one time, the brain is a complex system of neurons that excite or deactivate each other in large connected systems of neurons. These patterns of excitation and deactivation are the ways the brain processes, learners, and makes decisions.
The Unconscious System in humans is the source of intuition and of a good deal of human insight and creativity (McGilchrist, 2021). Intuition is the ability to understand something immediately without the need for conscious reasoning. An example of intuition is the common phenomenon where a scientist, mathematician, or artist has worked a great deal on a problem with conscious effort and only later, when they have stopped thinking about the problem, the solution to the problem just seems to come to them (Sfard, 1994).
The human web of associations represented by neural connections in the brain acts as a “prediction engine” (Eagleman, 2020; Seligman et al., 2016). The concept of the human brain as a prediction engine is a model in neuroscience that suggests that the primary function of the brain is to predict future events based on past experiences. As sensory data comes in, the brain compares the actual input with its predictions. If the prediction is accurate, the brain does not need to do much—it’s a confirmation that the model of the world it has built is correct. When there is a mismatch between prediction and actual sensory input, an error signal is generated. This error signal is crucial for learning and adaptation. It tells the brain that its model of the world is not accurate in some way, and adjustments need to be made. The brain then updates its predictions to account for the new information. This process is continuous and happens at various levels, from basic sensory experiences to complex social interactions.
The brain is Bayesian
Prediction is inherently a matter of computing probabilities. In fact, it likely the brain engages in Bayesian inferencing to establish its probabilities. Studies have shown that the brain integrates sensory information in a way that is consistent with Bayesian updating (Knill and Pouget, 2004). For example, when visual information is uncertain or ambiguous, the brain often relies on prior experiences (priors) to make sense of the sensory input, effectively making a “best guess” that aligns with Bayesian inference.
The brain’s ability to learn from experience and adapt to new situations also demonstrates principles akin to Bayesian probability. Through experience, the brain develops “priors” (probabilities that represent the degree of belief about a hypothesis before considering the current data or evidence) that inform future decisions. When faced with new experiences or information, these priors are updated, which is like how Bayesian inference incorporates new data to update beliefs.
Ways of meaning
Now, we will argue that Chat and humans deal with semantics in a similar way. Chat organizes “word” meanings in a multidimensional vector space where words closer in meaning are closer in the vector space. During training, the model learns to predict the next word based on its context (surrounding words) or to predict the context based on a word. Through this process, it captures semantic and syntactic regularities in language.
What Chat and other Large Language Models mean by “word” can vary. For Chat, at the simplest level, a word is treated as a sequence of characters bounded by spaces or punctuation in written language. However, in some AI systems, concepts and what they refer to in the world, not words, are represented in the semantic vector space. The key is that the geometric relationships between word or concept vectors in the space reflect the semantic relationships between the words or concepts themselves.
The brain takes in data through many more senses than Chat and has a much bigger neural system. Nonetheless, there are intriguing parallels and theories that suggest the brain may well represent concepts in a multidimensional vector space.
There is evidence to suggest that the brain represents concepts in a distributed manner across different neural networks (Zhang et al., 2020). This means that a single concept is not localized to one specific area of the brain but is instead represented by a pattern of activation across multiple regions. This is somewhat analogous to how word embeddings distribute the representation of a word across multiple dimensions in a vector.
Neuroimaging studies have shown that semantically similar concepts activate similar neural patterns (Patel et al., 2023). For example, when people think about tools or animals, the patterns of brain activity are more similar within each category than between categories, indicating a kind of semantic proximity in neural representation.
Just as word embeddings can be thought of as occupying points in a high-dimensional space, concepts in the brain might also be represented along multiple dimensions of meaning. These dimensions could correspond to different sensory, motor, emotional, and cognitive aspects of the concepts.
Categories (“Things”)
While visual images can be broken down and examined in granular terms in laboratories, human brains working out in the world perceive their surroundings not as an assemblage of colors and contours, but in terms of categories (things and processes) (Li, 2023). Categories (concepts when they in the mind) are the intersection between sensation (the way we sense things) and language (the way we describe them).
The human brain may well handle categorial meanings as a large vector space as do some AI systems. Part of the way categories are placed in this space is probably innate. Part is probably universal because it is based on the shared experiences all humans have. And part is diverse across people and cultures because of individual and cultural differences.
When an AI vision recognition device is trained to recognize categories, it is more important to show it lots of examples of what something is not than what it is. As Li (2023: 132) has pointed out in her groundbreaking work on AI vision: It was, ultimately, the diversity of other things our algorithm had seen that gave it a kind of perceptual experience and allowed it to perform so well when presented with something new. Rather than immersing the machine in hundreds of photos of airplanes covering as many variations of color, style, perspective, and lighting conditions as possible, we had shown it precisely one. We did, however, show it hundreds of images of completely unrelated subjects … So while it had been trained on a wide variety of things, the airplane it had just recognized was only the second it had seen. Ever.
This principle—that “the secret to recognizing anything was a training set that included everything”—is utterly Saussurean (Saussure, 1916).
The semantic categorial/conceptual vector space in our brains is an abstract ontology. When, in actual language use, people match words to this conceptual space for communication purposes they actively give these words additional and sometimes novel situational means based on their often-massive knowledge of context and their purposes. The semantic vector space fixes the range of possible situational means and this very fixedness is what allows for a wide range of contextual meanings, including new ones. New situational meanings may well change the shape of the vector space by creating new realizations about the relationship between concepts.
There is reason to believe grammar operates in much the same way as words, combining a fixed system and wide variation. Grammar in the brain is a model of the world at the logical propositional level (this was basically Wittgenstein and Ogden, 1922 point in the Tractatus). Units like “subject,” “object,” “clause,” and so forth, are categorial/conceptual fixed entities in a model of the world as sets of events. Each such unit can take on different situational meanings in use. For example, grammatical subjects can, in use—and in different languages—take on various specific situational meanings. For example the grammatical subject in English can take on these different meanings, and others, in contexts: Actor (“John left home”); Inanimate Agent (“The rock broke the window”); Experiencer (“She feels happy”); Patient (“The cat was chased by a dog”); Causative Agent (The professor made me angry”); Possessor (“My brother has a new friend”); Measure or Quantity (“Five miles is too far”); Description or State (“The girl is happy”); Phenomenon (“The flood happened late”); Existential (“There is a problem”); Impersonal (“It’s snowing”); Dummy Subject (It is necessary to pay attention”).
Situational meanings
Grammar is an abstract system that is used not only for communication but for thought as well. A word or image can be taken to represent a category (e.g., “bird”) or it can, in use, be associated with specific situational meanings created by speakers and hearers (Gee, 2017). Categories are not primarily about communication. They are about ontology. But once that ontology is set, we can “speak” in a wide variety of situationally diverse ways about that ontology. This, too is Saussurean; it is just Saussure’s distinction between langue and parole. There is no reason to assume that grammar (langue) is correlated one-to-one with function labels.
It is key, especially to readers of this journal, to see that situational meanings—the meanings of language in use—are fully multi-modal and multi-sensual not just verbal. If someone says to you “Eating durian is a novel experience,” if you only know that durian is a fruit all you have is categorical meanings and no real idea what “novel experience” means at a situational level. If you have tasted, seen, smelled, and read about durian, you bring an amalgam of different sensual experiences (sight, smell, taste, texture, memories, words read and heard) as the situational meaning for “durian” and can add depths of meaning to “a novel experience.”
“Durian” is not special. This is how all words work in use. If I say to you “America is a one-click democracy” you may well bring a wealth of lived experiences with the internet, shopping, “likes,” voting, disinformation, and elections to create a robust situational meaning for the phrase. This is why Borges (2000: 117) said: “Words are symbols for shared memories. If I use a word, then you should have some experience of what the word stands for. If not, the word means nothing to you.”
Meaning for language in use is experiential. Situational meanings are not words per se, but amalgams of elements of experience supplied by our senses, sensibilities, and various modes. Because people have different experiences in life, they bring different situational meanings to interaction and sometimes must negotiate them. When they negotiate, they are negotiating not what words mean, but what experiences mean.
Sensation and multimodality
In studies of multimodality, a “mode” refers to systems of communication that utilize different sensory aspects and cognitive skills. Just as Chat can use language quite well as communication (though not as a tool for thought), so, too, it and other Generative AI systems are getting better and better at using other modes and integrating them. It will eventually be quite good at multimodality. Multimodality is not where Chat departs strongly from humans. Sensation, feeling, and emotion is what Chat lacks and will never have until it has a body. By the way, the term “consciousness” can mean either “reasoning” or “being self-aware.” Chat can engage in reasoning, but it is not capable of awareness because awareness is a feeling.
Feelings and emotions are central to human thinking, deciding, and action. Our feelings and emotions—like hunger, excitement, caring, and fear—exist to help guide and assess action consciously (Barnett, 2017; Damasio, 2018; McGilchirst, 2021; Mlodinow, 2022; Solms, 2021). The feeling or emotion tells us that something is, here and now, going right or wrong with our body and that we should act to maintain, repair, or enhance the body’s homeostasis, its balance with the world. The feeling or emotion also assesses the action. The action is considered good if it maintains or enhances good feelings or emotions or if it lessens or removes bad ones. Furthermore, humans use feelings and emotions not only to assess the success of an action, but to affect how their brain stores, edits, and connects the results of experience.
Chat has no feelings or emotions and so, in principle, it is not like a human. And here no progress will be made for a long while. In this sense, Chat thinks and decides and communicates (a form of action) quite differently than humans.
Sensuous constructions
There is an aspect of meaning that is rarely discussed, though it is a form of multimodality—perhaps, better put, a form of “multi-sensuality/multi-experientiality.” This is what we call “sensuous constructions” (Zhang and Gee, 2023). Sensuous constructions are like situational meanings; they, too, are amalgams of sensation, feeling, emotion, sensibility, modes, words, and memories.
Anything—words, objects, categories, sensual experience (e.g., an image or a smell), etc.—can become sensuous constructions. Sensuous constructions are ensembles of feelings, emotions, memories, and embodied experiences we construct around something—in life or media—that we have experienced repeatedly across time and space.
For example, for “birders” birds or different species of birds (as a category or a word) become quite different and more expansive sensuous constructions than they are for non-birders. “Pileated Woodpecker” is a category and word, but for a given birder it is whole spatiotemporal pattern/gestalt of merged sensations, feelings, emotions, texts, communications, interactions, and experiences. It is this lived reality over time and space that imbues human life with its deepest meanings and values. People’s bodies as they move through space and time create sensuous constructions out of things, words, images, and other sensory experiences.
The 20,000-year-old majestic image of a dying bison on a wall in Lascaux Cave in southwestern France was incorporated into spiritual rituals and its major purpose was to point to a sensuous construction (Zhang and Gee, 2023). It was meant to capture not an individual bison and not just the idea of “bison species,” but the majesty, power, and spiritual meaning of bison as (at the time) beings dominate to humans. It was composed of sensuous details and all the experiences these people had had with bison and nature and it made a certain inner sense of the bison. The ancient humans who entered this cave carried with them into this cave, through their lived experiences, a sensuously, emotionally, and spiritually replete “concept” of bison that was neither a word, a category, or a conversationally specific situational meaning.
Sensuous constructions do not represent anything, they are more like stories than references. In the famous anime Attack on Titan a red scarf becomes an object around which multiple embodied experiential meanings cluster, created in many different scenes across the whole series. Words, things, images, places, events, smells (remember Proust’s petites madeleines) and other sensations, can become a point around which a myriad of multisensory, multimodal, affective components of embodied experience cluster, changing with each new experience.
Chat has had no embodied experiences over time and so cannot form sensuous constructions. Yet these are crucial to our deepest sense making as humans. Sensuous constructions have been little studied. Since Chat will eventually be pretty good at multimodality, words, and referential meanings, students of multimodality may want to turn to the study of how feelings, emotions, sensations, and embodiment meld to create situational meanings and sensuous constructions.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
