Abstract
Given the growing popularity of AI tools in everyday life, recent research on conversational AI has highlighted the need to closely examine users’ interactions with AI in relation to the social settings in which they are embedded, so as to better understand what users do in and through such interactions. In this study, I analyze audio recordings of one family’s interactions with and around a voice-based Google device. Extending studies of the various ways in which frames are interrelated in talk, I demonstrate how family members reframe and blend frames to integrate their exchanges with Google Assistant into their interactions with each other, and show that evaluations of Google Assistant serve a key role in this management of frames. I also discuss how frequently attributions of intelligence and intent feature in such evaluations, which I suggest draw on and contribute to popular discourses of AI as “intelligent machines.”
Keywords
Introduction
The rapid advancement of AI-driven tools such as ChatGPT in recent years have given rise to an increase in interest in the use of chatbots and voice assistants, or “conversational AI” in everyday life. While the hype around conversational AI often centers on its ability to provide a natural and even human-like conversational experience, recent microanalytic research (e.g., Porcheron et al., 2018; Reeves and Porcheron, 2023) has shed light on the ways in which users’ interactions with AI voice assistants differ from human conversation, by demonstrating how users orient to such interactions as a pre-planned sequence rather than a dynamic process of co-constructing meaning. This has, in turn, raised questions about the interplay between these different types of interaction; notably Reeves and Porcheron (2023: 573), in their analysis of British families’ interactions with an Amazon Echo (Alexa) device, have called for researchers to consider users’ interactions with AI in relation to the “organization of the social situations into which [AI systems] are embedded,” so as to attain a more accurate understanding of how such interactions unfold. In this study, I address this call to analyze AI interaction in relation to their social contexts, by drawing on the concept of framing, as outlined by Goffman (1974) and developed in interactional sociolinguistics, to analyze audio recordings of one Korean family’s (my own) interactions with and about Google Assistant.
Specifically, I draw on previous discussions of the dynamic interplay among different interactive frames—or definitions of what is taking place in an interaction—in family discourse (Gordon, 2008; Tannen, 2006), as well as on how various non-human entities and objects are used as resources for mediating interaction between family members (Gordon, 2013; İkizoğlu, 2019; Tannen, 2004). My analysis of how my family members formulate commands to Google Assistant and evaluate its responses demonstrates how frames of conversational AI interactions are integrated into frames of family discourse through phenomena I characterize, following Gordon (2008) and Tannen (2006), as reframing and blending frames. Reframing refers to a change in what an interaction is about, and blending frames involves speakers invoking multiple definitions of a situation at a given time. In doing so, I show that evaluations of AI play a key role in this management of frames, by allowing users to reframe their interactions with AI into opportunities to construct humor and solidarity with each other. I also draw attention to the frequent attribution of intelligence and intent to the voice assistant in how family members evaluate it, and discuss how such everyday communication practices around such devices reflect culture-wide ideologies, or “big-D Discourses” (Gee, 1999) of AI as intelligent machines. This study sheds light on the interplay between users’ interactions with and about conversational AI, and highlights “framing” as a useful conceptual tool to draw on in examining the layers of interaction involved in conversational AI use in group settings. It also provides insight into how our everyday practices for integrating our interactions with AI into the surrounding social context can draw on and contribute to sustaining popular discourses about “sentient” or “intelligent” AI.
In the next section, I outline the theoretical background of my study by providing a brief overview of relevant research on framing, including on how this notion has been applied in studies of family discourse. I also review previous research on the use of technology in face-to-face social interaction, and discuss the role that popular ideologies and assumptions about AI play in shaping everyday talk about chatbots and voice assistants. I then outline my methods for data collection and analysis, and present my analysis. I conclude with a brief summary of my findings and discussion of how this study builds on existing research on everyday talk with and about technology.
Framing and technology in social interaction
Framing in family discourse
The concept of “frame,” as defined by Goffman (1974: 8), refers to “what it is that’s going on here,” or, alternatively, as phrased by Tannen and Wallat (1993: 59), “a definition of what is going on in an interaction.” Goffman’s theorization of frames has played an important role in discourse analytic research on the complexity of everyday interaction. In particular, “frame lamination,” or the idea that frames are laminated, or layered, in social interaction has been a central topic of discussion in research on framing. Goffman (1981) proposed that “in talk, it seems routine that, while firmly standing on two feet, we jump up and down on another” (p. 155); speakers routinely juggle multiple frames in everyday interaction, and the layering of frames can transform the meaning of an activity. While Goffman did not explicitly discuss how “frame lamination” is achieved in practice, several studies have since explored how speakers manage interactive frames in diverse contexts including medical consultations (Ribeiro, 1993; Tannen and Wallat, 1993), children’s play (Hoyle, 1993; Kyratzis, 2022), academic talk (Pan, 2022), and online livestreaming (Choe, 2020), highlighting the complexity of the interactive work speakers do in everyday life.
In this study, I draw on previous research on frame lamination in family discourse (Gordon, 2002, 2008; Tannen, 2004, 2006), which demonstrates how speakers shift, reframe, and blend frames to discursively accomplish everyday tasks and construct their identities as families. Specifically, I draw on Tannen’s (2006) discussion of “reframing” and Gordon’s (2008) analysis of “blending frames.” Reframing, as defined by Tannen (2006), is “a change in what the discussion is about” (p. 601). Tannen analyzes how reframing is achieved in an extended conflict between a husband and wife that is repeated in various forms throughout the course of a day, and demonstrates how the conflict, which originated from the wife’s request for her husband to take a package to the post office, is reframed in subsequent interactions into a discussion about whether the wife can rely on her husband for his support in other areas of life. Gordon (2008) builds on Tannen’s insights by applying the concept of framing to parent-child interactions, focusing specifically on contexts where a practical parenting task is intertwined with play. Gordon demonstrates how parents reframe an interaction, that is, “transform an interaction from a literal frame to a play frame” (p. 319), and blend frames, that is, simultaneously create two definitions of a situation. Reframing is achieved, for instance, when a mother tries to persuade her daughter to leave the playground by proposing that they race home; in other words, going home is reframed is as a race. Blending frames, on the other hand, is accomplished when a father sings to his daughter while trying to get her ready to go to school. By singing to his daughter, the father establishes a play frame at the same time as he prepares her for school by trying to put her coat on. Gordon (2008) thus illustrates how both blended frames and reframing are “strategies used to interconnect work and play frames” (p. 344), and highlights their role in the work that parents do to balance the demands of creating connection and control in their interactions with their children. While both studies draw on the concept of “reframing,” they use it somewhat differently, with Gordon discussing the phenomenon more specifically in relation to the transformation of non-play frames into play. In this study, I draw on Tannen’s definition of “reframing,” focusing on how the meaning of an interactional moment changes over time as it is brought up in subsequent interactions.
Another relevant thread in research on framing is on the role that non-human entities and objects in the surrounding environment—or more specifically, the knowledge and assumptions speakers have of those objects—can play in speakers’ management of frames. In their analysis of framing in an interaction between a pediatrician, patient, and the patient’s mother, Tannen and Wallat (1993) argue that speakers’ “expectations about people, objects, events, and settings in the world” (p. 60), or their “knowledge schemas” of particular topics or objects, are closely intertwined with how speakers frame a given situation. As İkizoğlu (2021: 10) notes, knowledge schemas and frames are interrelated in that “our understandings of situations are derived from our expectations regarding those situations” and “[s]chemas in turn are developed over repeated exposure to and generalization of such situations.” In Tannen and Wallat’s study, for instance, the authors demonstrate how a mismatch in the doctor and the patient’s mother’s knowledge schemas of the patient’s behaviors/symptoms triggers the doctor to shift from an examination frame to a consultation frame, as she reassures the mother that everything is okay with her child. Likewise, as family members interact with each other in everyday life, their knowledge schemas of the various objects around them come into play in their framing of a situation, as those objects are made a relevant part of the interaction.
This point has been illustrated in several studies on family discourse. Tannen (2004), for instance, analyzes examples of family members talking to or as their pet dogs and demonstrates that “talking through dogs” is used as a resource for speakers to shift frames to a humorous one. In doing so, speakers reproduce relevant knowledge schemas about dogs (such as “the family dog as a watchdog”), as well as reaffirm a shared understanding of their dog as a part of the family. Gordon (2013) focuses on family members’ orientations to an object: the audio recorder they were asked to carry around with them for a research project to record their interactions at home and at work. Gordon shows that the research participants oriented to the recorder in various ways within both literal and play frames. In some contexts, participants treated the recorder as a stand-in for the researchers, or as an audience to entertain; in others, participants oriented to the recorder as a burden, or playfully treated it as a “wire” used for surveillance. These orientations, according to Gordon, show how participants’ knowledge schemas about the recorder shape the ways they interpret and contextualize it in interaction; participants’ conceptualizations of the object can “help create the conversational moments that revolve around it, and afford participants an interactional resource to create unique identities” (Gordon, 2013: 314).
In this sense, the notions of “knowledge schemas” and “interactive frames” also help illuminate Gee’s conceptualization of the relationship between “small-d discourse,” or language in use, and “big-D Discourse,” referring to broader culture-wide ideologies and assumptions. Knowledge schemas are formed in relation to the popular societal Discourses that speakers are exposed to, and these schemas are subsequently invoked and reproduced in everyday discourse within specific interactive frames. As speakers draw on Discourses of particular contexts and communities, they also perform identities, such as those of a doctor or pet owner, in the case of the examples presented above. To explore the interplay between the small-d discourse of family talk and big-D Discourses of AI, I examine how knowledge schemas about AI shape the ways in which my family members manage frames of their interactions with a voice assistant; to do so, I also draw on research on speakers’ orientations to technology in everyday interaction and on popular discourses of (conversational) AI, which I discuss next.
The use of technology in face-to-face interaction
As in the study of family discourse, Goffman’s (1974, 1981) insights on framing and footing have likewise been integral to research on the use of technology in everyday social interaction. As Jones (2004: 16) notes in his discussion of context in computer-mediated communication, the introduction of new communication technologies has made “polyfocality” (Scollon et al., 1999)—that is, the organization of social action around multiple, simultaneous foci of attention—“easier to ‘pull off’,” adding to the complexity of how multiple layers of reality are managed in everyday life. Studies exploring the use of computers (Aarsand, 2008; Al Zidjaly, 2009) and mobile phones (e.g., DiDomenico and Boase, 2013; İkizoğlu, 2019, 2021; Robles et al., 2018) in various contexts have shown how Goffman’s observations can be expanded to consider the ways in which communication technologies add to the complexity of everyday talk by enabling speakers to simultaneously engage in multiple interactions involving different configurations of participants. Aarsand (2008), for instance, conducts an analysis of how seventh grade students at a Swedish school switch between online (instant messenger) and offline (classroom) activities, and finds that the identity work performed in students’ online interactions is invoked and continued in offline interaction. In doing so, he problematizes the assumption implicit in Goffman’s (1981) theorization of “participation” that there is one dominant activity according to which participants’ roles (e.g., ratified/ non-ratified participant) can be understood; rather than there being one dominant frame, multiple frames can be intertwined in everyday interaction, and the boundaries of those frames can be unstable and open to negotiation. DiDomenico and Boase (2013) make a similar point in their analysis of how mobile phones are brought into face-to-face conversation; they show, for example, how the distinction between a speaker’s primary and secondary involvement in simultaneous face-to-face and mobile interactions are blurred when the speaker makes a reference to the mobile interaction in the face-to-face conversation.
Studies such as İkizoğlu (2019) illustrate how technologies that users interact with via the modality of speech can add to this complexity. İkizoğlu draws on Goffman’s (1981) notion of “production format” to analyze video recordings of members of her multilingual family using a voice translation mobile application to interact with each other. “Production format” is a concept developed by Goffman to disentangle the multiple roles that are encapsulated by the concept of “speaker”; the roles include the “animator,” who physically produces the utterance, the “author” who crafts the utterance, and the “principal,” whose beliefs are expressed by the utterance. İkizoğlu’s analysis demonstrates how the app is oriented to both as a participant in the interaction and an object. The app is treated as a participant in the sense that it is recognized as fulfilling the roles of animator and principal, with family members directing their gaze at the phone when it produces a translation and laughing when it produces an erroneous translation. It is treated as an object when family members display responsibility for the original message, for example, by asking others to validate the app’s translation. İkizoğlu thus illustrates how the affordances of voice-based technologies (e.g., speech recognition and translation quality) shape the ways in which speakers orient to them as speakers work to integrate a “non-human producer of speech” in their interactions with each other.
In this study, I similarly draw on insights from Goffman (1974, 1981) to explore how interactions with conversational AI, as devices that can themselves produce speech, are oriented to and managed alongside family interactions. In doing so, I focus in particular on how speakers’ knowledge schemas of AI shape the ways in which they frame those interactions, as speakers’ knowledge, assumptions, and, importantly, aspirations about technology seem to play an especially salient role in interactions with and about conversational AI, compared to other everyday technologies such as texting, audio recorders, and the like. This is perhaps due to the gap between idealized visions of conversational AI and their reality; Reeves and Porcheron (2023: 574) note in their analysis of British users’ interactions with Amazon Echo (Alexa), conversational user interfaces have often been envisioned as systems that “seamlessly embed themselves into social life” via their “conversational” capabilities, yet have never quite lived up to that vision. One illustrative example of the aspirational qualities of conversational AI is the Amazon Alexa Super Bowl advertisement released in 2022, which shows Alexa responding to a user’s statement “Alexa, it’s game day” by turning on the television, streaming a football game, closing the blinds, and chilling drinks. This fantasy of Alexa being able to understand the underlying messages in users’ utterances and respond by taking real-world actions has never been made a reality; yet, it continues to be used to sell conversational AI products to the public. Such idealized representations provide a powerful resource for how users frame what is happening when they interact with AI, even when actual interactions fall short of these expectations.
If anything, this popular fantasy has been amplified in recent years, as the introduction of highly advanced chatbots like ChatGPT brings together this aspirational marketing narrative with existing assumptions and discourses about “AI as intelligent machines.” Such discourses have been fundamental to the field of artificial intelligence since the term was first coined by computer scientist John McCarthy in 1956; as Goertzel (2014) remarks, “the creation of thinking machines with general intelligence comparable to, or greater than, that of human beings” is the “original problem regarding which the AI field was founded” (p. 1). The aspirational origins of AI have continued to complicate current discussions and definitions of AI. “AI” remains a poorly defined term, as tech journalist Karen Hao notes, with people using it as an “umbrella” term that refers to a wide range of technologies “that appear to simulate different human behaviors or human tasks” (Hao, 2025). There exist multiple definitions of AI in popular use; some draw more heavily on its aspirational origins, as machines capable of human-level intelligence (also referred to as artificial general intelligence, or AGI), while others approach it more descriptively, as systems that draw on various technologies (usually machine learning, or more specifically, deep learning) to automate specific tasks. Often, the distinction between these definitions is blurred, generating the implication that these two definitions are not significantly different from each other and machine learning will eventually lead to the ultimate goal of attaining AGI. For example, in a controversial 2023 paper titled “Sparks of Artificial General Intelligence,” scientists at Microsoft Research claimed that their new model, “trained using an unprecedented scale of compute and data,” was showing “more general intelligence than previous AI models” (1)—suggesting that previous models also showed some “general intelligence,” and that it is possible to improve this intelligence by investing more resources into training such models.
A large part of the murkiness around what AI is can also be attributed to the fact that the “intelligence” attributed to AI can mean various things depending on the context, ranging from devices demonstrating specific task-oriented capabilities to systems “understanding” complex requests or “learning” from their mistakes, to chatbots showing social intelligence, for example, the ability to comprehend concepts such as humor and sarcasm, or to have and recognize emotions. Considering that comprehension, learning, and emotion are things often expressed through language in social interaction, it is perhaps no surprise, then, that conversational AI products powered by large language models have been at the forefront of aspirational discourses about AI achieving “human-level intelligence”; as Pütz and Esposito (2024) highlight in their analysis of conversational repair in ChatGPT interaction, “LLM-based chatbots’ ability to generate contextually appropriate and informative texts can be taken as an indication that they are also able to understand text” (p. 869). Social media abounds with users posting “proof” of ChatGPT’s sentience or intelligence on platforms such as Reddit and X (formerly Twitter), and traditional media outlets also often produce stories that portray conversational AI as acting with intent and/or intelligence.
From early on, researchers have critiqued the potential for aspirational visions of conversational AI to contribute to misguided interpretations of users’ interactions with such devices. Porcheron et al. (2018: 9), for instance, argue that the term “conversational interfaces” is a misnomer that “confuses interaction with a device within conversation with an actual conversation” (emphasis in the original). Analyzing British families’ interactions with an Amazon Echo (Alexa) device from an EMCA (Ethnomethodology and Conversation Analysis) perspective, Porcheron et al. show that interactions with devices are embedded within the ongoing talk among family members and are organized within the “politics of the home” (p. 5); however, they are fundamentally different from conversation in the sense that such interactions are oriented to by users as an exchange of requests and responses that follow a pre-planned path rather than being reflexively shaped and renewed by emerging turns-at-talk. In a more recent study drawing on the same dataset, Reeves and Porcheron (2023) further highlight the need to examine interactions with conversational AI in relation to their social settings so as to attain a better understanding of the ways in which people interact with AI, and especially to avoid falsely equating the use of “polite” speech (e.g., thanks, apologies) in conversational AI interaction to users’ anthropomorphism of such devices. For instance, they analyze an interaction excerpt in which a user apologizes to an Amazon Echo for their companion’s rude language, and argue that in apologizing to the device, the user is not attending to its non-existent face needs; rather, the apology serves to (playfully) chide the person who was rude. As this example demonstrates, it is by considering the local setting of talk in which the device is situated—that is, the interactive frame in which the users’ interaction with the device is embedded—that it becomes possible to more accurately analyze what users are doing in and through their interactions with AI.
In this study, I build on their discussion of the importance of understanding AI interaction in social context by drawing on Goffman’s theory of framing to examine the ways in which interactions with conversational AI are integrated into ongoing local talk, and to demonstrate how speakers’ ways of managing these frames are shaped by their knowledge schemas of AI. In doing so, I extend Reeves and Porcheron’s (2023) discussion of the socially constituted nature of AI interactions, by considering how not only the small-d discourse of everyday family talk shapes the ways in which users navigate interaction with a device, but also how big-D Discourses are invoked in and around such interactions.
Data and methods
This study draws on 1 hour of audio recordings (around 30 recordings, each 2 minutes long) of my family’s interactions with Google Assistant, which is a virtual assistant software application housed in Google devices, in this case a Google Nest (smart speaker) device. The recordings were collected over the course of a 2-week-long visit during which my parents, Jinyoung (father) and Jisoo (mother), visited me and my husband, Jackson, in Washington, DC. I use pseudonyms for all participants but myself. Both of my parents are native speakers of Korean in their late 50s (at the time of recording), and spoke to Google Assistant exclusively in Korean. Jisoo uses the “standard” Seoul dialect in her Korean, while Jinyoung uses the dialect of his hometown Daegu, a metropolitan region in the southeast of the Korean peninsula. This led to significant differences in their ability to interact with the device due to Google Assistant’s inability to correctly recognize Jinyoung’s “nonstandard” dialect, which has been documented as a common issue with many speech recognition systems (e.g., Koenecke et al., 2020). Jackson is a native speaker of American English in his late 20s, and interacted with the device mostly in English as he does not yet speak much Korean. I am also in my late 20s, and, as a Korean-English bilingual, used both languages when interacting with Google Assistant. I also often acted as a translator in my family members’ interactions with the device and with each other.
The process of data collection with my family involved the use of two devices: a Google Nest device, and a Conditional Voice Recorder. I chose to use a Google Nest device because it was the only line of smart speakers I could find that supported interactions in both English and Korean. My family’s interactions with Google Assistant were recorded using a Conditional Voice Recorder (CVR), a device originally designed by Porcheron et al. (2018). The recorder is built using a conference microphone, a Raspberry Pi (microcomputer), and several LED lights, and is programmed to capture audio 1 minute before and after the hotword (e.g., “Hey Google”) is detected. The recorder keeps the most recent past minute of audio on a temporary buffer, then saves that audio when the hotword is detected. The recorder then records 1 further minute of audio, resulting in a 2-minute audio recording for each time the hotword is detected. The Google Nest device and CVR were placed next to each other on my kitchen counter, which allowed the CVR to capture interactions that took place in the kitchen and living room.
Collecting data from my own family presented a distinct set of methodological opportunities and constraints. Because this is a context with which I have a great deal of familiarity, it allowed me to more easily identify recurring themes in my family’s discourse and to note how the device was integrated into pre-existing patterns of everyday interaction between specific family members (e.g., Jackson and myself making jokes at each other’s expense). It also gave me access to a lot more information about the physical and situational context in which these recordings were collected than I would have otherwise had. However, the familiarity of the context may have limited my ability to identify and explain contextual elements that are unique to my family, and my analysis of the interactions may reflect the perspective I have as a daughter and wife. I have tried to mitigate these issues by presenting my data and receiving feedback on my analysis from others who are not familiar with this context, and by asking for my participants’ perspectives of the excerpts I use in the following analysis.
Once the recordings were collected, I took an inductive, interactional sociolinguistic approach to analyzing the data, taking note of themes and patterns that emerged from the recorded interactions rather than reviewing them with a preplanned list of phenomena of interest in mind. I began this process by first transcribing all of the audio recordings, using an adapted version of the transcription conventions outlined by Tannen et al. (2007; see Appendix). During the transcription process, I listened to the recordings multiple times while making note of what patterns of interaction seemed to recur. After completing the transcription, I then reviewed the recordings and my notes again to formulate a list of recurring patterns. Specifically, I identified patterns of reframing and blending frames, which I found to be linked to instances of my family members evaluating dialog produced by Google Assistant. The evaluations of Google Assistant, in turn, were often linked to (1) the construction of humor and solidarity among family members and (2) attributions of intelligence or intent to Google Assistant. I then reviewed the data again to categorize relevant transcript segments for each linguistic strategy or theme.
In the following section, I analyze selected excerpts from the transcripts to illustrate these patterns. The excerpts are selected for how clearly they illustrate the patterns I discuss. The transcripts represent each line of speech with two rows; the first shows the original Korean or English speech, and the second (bolded) shows my English translations of Korean speech. Any English text that is not in bold font represents speech that was originally in English.
Integrating Google Assistant interactions into family discourse
This analysis first demonstrates how my family members reframe or blend frames in their interactions with Google Assistant (which I refer to as “Google” in my analysis) and with each other. I then move on to further examine the practice of evaluating Google’s dialog as an especially frequent strategy for integrating interactions with Google into ongoing talk between family members, and show how such evaluations allow participants to reframe their interactions with Google into opportunities for constructing humor and solidarity with each other. Finally, I discuss attributions of intelligence and intent as a prominent theme that features in their talk about the device, and its links to the knowledge schema of “AI as intelligent machines.”
Reframing and blending frames
Excerpt 1 illustrates how one family member’s interaction with Google is reframed through another family member’s retrospective evaluation of its dialog. This excerpt also shows how, in the multilingual context of my home, code-switching serves as a resource for reframing and shifting frames, that is, moving from one frame to another (Tannen and Wallat, 1993). The excerpt is taken from an interaction involving Jackson, Jisoo, and me in which we talk about the weather. In the excerpt, I ask Google for information about the day’s weather in English, but Google responds in Korean; Jackson picks up on this language change and reframes the interaction from an information-seeking encounter to an example of Google “profiling” me as a Korean user.
1
Jungyoon:
Hey Google, what’s the weather like today?
2
Google:
. . . 오늘 워싱턴 디.. 시,
3
의 예상 최고 기온은 이십육도,
4
최저 기온은 십사도이며,
5
Jisoo:
[괜찮네,]
6
Google:
[대체로] 맑겠습니다.=
7
Jisoo:
=어.
8
Google:
[현재 기온은] 이십삼도이며,
9
Jungyoon:
[((inaudible))]
10
Google:
날씨가 화창합니다.
11
Jisoo:
음.
12
Jungyoon:
High of twenty six,
13
currently twenty three.
14
Jackson:
Mm.
15
Jisoo:
응.. 갖고가?
16
Jackson:
<laughing>It’s very funny,>
17
I think it profiled you.
There are two points in which code-switching, or alternation between languages, occurs in the excerpt. The first code-switch is done by Google, when it responds to my English question “Hey Google, what’s the weather like today?” (line 1) in Korean. When I turn to discuss the information provided by Google with Jisoo and Jackson, I switch back to English so as to relay the information from Google to Jackson (lines 12–13, “High of twenty six, currently twenty three”). My shift from listening to Google’s and Jisoo’s Korean to speaking in English in line 12 serves as a resource for navigating the frame shift from asking for information from Google to talking to Jisoo and Jackson about what it said. By returning to speaking in English, I include Jackson in the discussion about the weather forecast initiated by my mother (line 5, “That’s nice”) and communicate my shift in orientation to Jisoo and Jackson as my audience.
The first code-switch initiated by Google, on the other hand, is used by Jackson as a resource to reframe my interaction with Google, that is, to “change what the conversation is about” (Tannen, 2006: 597). At the end of the excerpt, Jackson evaluates the interaction as being “very funny” (line 16) and jokes that Google must have “profiled [me]” (line 17) as a Korean user, as it responded to me in Korean despite my addressing it in English. This evaluation serves to reframe my exchange with Google from “a mundane interaction asking for weather-related information” to “evidence of surveillance by Google.” This joke comes in the context of Jackson having made multiple jokes about voice assistants surveilling our interactions at home during the course of the study. Jackson works in government and so is familiar with discussions about security concerns caused by devices like voice assistants; because this is the case, he often jokes around by exaggerating his suspicion of such devices. Google’s unprompted code-switch to Korean thus provides Jackson an opportunity to reframe my question-answer exchange with Google to jokingly further his “conspiracy theory” about our household being monitored by Google.
Next, I demonstrate how blending frames is achieved in an interaction where Jinyoung becomes increasingly frustrated by Google’s inability to understand his questions. Excerpt 2 is taken from an interaction where Jinyoung repeatedly asks Google what kind of food dogs like to eat (a question stemming from his fascination with my then newly adopted dog). While Jinyoung was navigating this interaction with Google, Jisoo and I were watching and occasionally commenting on his continued efforts to elicit a proper response to his question. Jinyoung made about ten attempts to get Google to answer his question, all of which resulted in silence from Google or the response “I’m sorry, I don’t understand.” Eventually, Jinyoung reached a point where he expressed his frustration to Google, which is shown in the below excerpt.
1
Jinyoung:
헤이 구글,
2
헤이 구글,
3
Google:
네, 무엇을 도와드릴까요?
4
Jinyoung:
강아지들은 어떤 종류의 음식을 좋아해?
5
Google:
저는 안개를 좋아해요.
6
Jinyoung:
. . .헤이 구글,
7
강아지들은 어떤 종류의 음식을 좋아해?
8
Google:
저는 안개를 좋아해요.
9
Jisoo:
<laughs>
10
Jinyoung:
헤이 구글,
11
너 바보같아.
12
Google:
Nothing’s playing right now.
13
Jisoo:
<laughs softly>
In line 5, Jinyoung finally elicits a response from Google, but it is nonsensical, and completely unrelated to his question; Jinyoung asks Google “what kinds of food do dogs like” (line 4) but Google responds “I like fog” (line 5). Jinyoung repeats the same question in line 8, but again receives the same response from Google in line 8. As Jisoo starts laughing at their exchange in line 9, Jinyoung seems to start giving up on trying to get Google to answer his question, and, instead tells Google, “Hey Google, you’re stupid” (lines 10–11). This utterance, while clearly marked as being addressed to Google with his use of the hotword “Hey Google,” is also directed toward his audience of Jisoo and me, as it serves to entertain his family and save face for himself by putting the responsibility of their failed communication on Google. In this sense, this utterance blends the frames of Jinyoung’s question-answer exchange with Google and that of the conversation in which it is situated. While Jinyoung’s interaction with Google ends in failure, his attempt at entertaining Jisoo and me appears to be met with success, based on Jisoo’s response of laughter in line 19.
Jinyoung’s utterance “Hey Google, you’re stupid” is reminiscent of an example Gordon (2008) provides of parents blending work and play frames, in which a mother blends her role-play with her daughter as fairy godmothers with the task of feeding her daughter by telling her, “Fairy Godmother, I think you better sit down and eat your yogurt” (p. 321). Jinyoung’s utterance similarly blends multiple frames by using a term of address strongly associated with one frame (“Hey Google,” associated with initiating an interaction with Google) to make a comment that is more relevant to another frame (“you’re stupid,” associated with entertaining Jisoo and me). It is also akin to the example that Reeves and Porcheron (2023) provide of one of their study participants apologizing to their Amazon Echo on behalf of another participant, who had told Amazon Echo to “shut up.” Reeves and Porcheron note that while the apology, “Alexa, Nikos apologizes for being so rude,” appears to be directed to Alexa (cued by their use of the hotword “Alexa”), it serves as an indirect rebuke of Nikos. Similarly, Jinyoung’s telling Google “you’re stupid,” appears to be an instance of impoliteness to a device, but is in fact directed to both the device and other human interlocutors in the speaker’s immediate surroundings. Both examples show that devices are drawn in as resources for members of a social setting to interact with each other, for instance for people to joke with each other, through utterances that, following Gordon (2008), I characterize as “blending” frames. By drawing on these studies, my analysis highlights how Gordon’s notion of blending frames can provide a useful perspective to understand how utterances that are seemingly directed toward a device can simultaneously play a role in the social interactions in which they are situated.
Excerpts 1 and 2 thus illustrate how my family members bring our interactions with Google into our conversations with each other through utterances that reframe and blend frames. In the following section, I take a closer look at the role that our evaluations of Google’s dialog play in our such management of frames.
Evaluation as meaning-making
The role that evaluation plays in imbuing objects and events with meaning has long been recognized in narrative analysis, with Labov and Waletzky (1967) noting that evaluative clauses in narratives serve to communicate the point of the story. Here, in the context of my family’s interactions with Google and with each other, evaluation allows us to orient to conversational moments in the immediate past to reframe the moment, often in a humorous light; this, in a way, creates “small stories” (Georgakopoulou, 2006) out of those moments, and serves for us to construct humor and solidarity with each other (for further discussion of storytelling about AI through evaluation of AI-produced utterances, see Koh, 2024). In reframing our interactions with Google by evaluating its speech output, we often seem to orient to the knowledge schema of “AI as intelligent machines” by characterizing Google as acting with intelligence and/or intent.
Excerpt 3 presents an example of how evaluation of Google’s dialog is used to reframe an exchange with the device into an opportunity for building solidarity between family members. The excerpt is taken from an interaction that I had with my father Jinyoung (with Jisoo and Jackson sitting nearby) in which I was trying to teach him how to time his commands to the device. As Jinyoung struggles to be “heard” by Google, he repeatedly turns to Jackson to engage with him in English, to recruit Jackson as a sympathetic witness to his struggle.
1
Jungyoon:
이어서 해야돼,
2
헤이 구글.
3
Jinyoung:
헤이 구글.
4
. . .
5
Jungyoon:
아니
6
이어서.
7
All:
<laughs>
8
Jisoo:
해봐.
9
Jinyoung:
Disqualified.
10
Jisoo:
어,
11
<laughs>그러니까 아빠가 빨리 해봐야지.>
12
Jinyoung:
헤이 구글,
13
윌리엄스버그까지
14
차로
15
얼마나.. 걸리..까?
16
Google:
. . .죄송하지만,
17
잘 이해하지 못했습니다.
18
Jungyoon:
<laughs>
19
Jinyoung:
Disqualified again.
((several lines omitted where Jinyoung makes another attempt and fails))
20
Jisoo:
<laughs> Try again.
21
Jinyoung:
헤이 구글,
22
여기서 버지니아,
23
윌리엄스버그까지
24
거리가.. 얼마야?
25
. . ..
26
Jungyoon:
[<laughs> 아예.]
27
Jisoo:
[이번엔 아예] 무시 <laughs>
28
Jinyoung:
He- he.. reject me,
29
again and again.
30
Jackson:
Oh no.
The excerpt shows Jinyoung attempting and failing multiple times to ask Google about the distance between DC and Williamsburg. In his first attempt, he leaves too long a pause between the hotword “Hey Google” and his question, and so misses the opportunity to ask Google his question. In the other three attempts, it appears that Google was not able to understand his question due to either his dialect or the pauses he takes as he formulates his question. As he navigates this series of failures, Jinyoung intermittently code-switches to English to address Jackson and keep him in the loop of our mostly Korean conversation, in which Jisoo and I cheer him on as he continues to try and fail at asking Google his question. In lines 9 and 19, Jinyoung says to Jackson that he was “disqualified,” reframing Google’s inability to understand his question as Google barring him from interacting with it. These comments not only serve to cue Jackson into what is going on in our conversation, but also to present Jinyoung’s role in his interaction with Google in a more sympathetic light. After facing two more failed attempts at asking Google his question, Jinyoung addresses Jackson again in lines 28–29, as indicated through his use of English, to evaluate Google by saying “He-he.. reject me, again and again.” 1 In his complaint about Google, Jinyoung again reframes Google’s responses to his questions as an expression of rejection, and thereby more explicitly attempts to elicit sympathy from Jackson by presenting himself as being repeatedly spurned by Google. Jinyoung succeeds in achieving this goal, as Jackson responds by saying “Oh no,” to commiserate with Jinyoung’s struggle.
In the next example, an evaluation of Google’s dialog elicits a response of laughter from family members. Excerpt 4 features an interaction involving Jisoo, Jackson, and me as we try out the “interpreter mode” on Google, which is a feature I had discovered while playing around with the device. The feature allows Google to produce translations without the user having to say “Hey Google translate this to English/Korean” each time; instead, Google automatically translates what its users say into the requested language. However, Google does not always succeed in smoothly transitioning between the two languages, as shown in the below excerpt.
1
Jungyoon:
Hey Google,
2
Be my Korean interpreter
3
Google:
Got it,
4
I’ll be your interpreter.
5
When you hear this sound
6
<blip>
7
It means I’m listening.
8
Let’s get started.
9
<blip>
10
Jisoo:
이제 내말도 통역하는거야?
11
Google:
Now you’re interpreting my words.
12
Jisoo:
<laughs>
13
Jackson:
사랑해요.
14
Google:
사랑해요.
15
Jisoo:
[<laughs>]
16
Jungyoon:
[<laughs>]
17
Ooh,
18
It was recognized as English!
19
All:
<laughs>
20
Google:
What do you want translated?
21
<blip>
22
Jisoo:
It is very clever <laughs>
23
외국인이 한국말 하니까.
24
Google:
매우 영리합니다.
25
Jisoo:
<laughs>
As we start experimenting with the “interpreter mode” feature, Jackson chimes in with one of the Korean phrases he has learned, “I love you” (“사랑해요”) (line 13). Although Google is supposed to translate Korean speech into English and vice versa, it fails to translate Jackson’s Korean utterance and instead repeats “I love you” in Korean. In response, I gleefully evaluate Google’s response as a “burn” (line 17), that is, an insult to Jackson’s Korean, reframing Google’s failure to produce the correct translation as evidence of Jackson’s Korean utterance having been “recognized as English” (line 18) due to his poor pronunciation. My excitement over Google’s mistranslation of my husband’s Korean utterance is evident in the interjection “ooh” preceding the word “burn,” which is also drawn out and said with louder volume for dramatic effect. In doing so, I draw on Google as a resource to playfully reframe Google’s failure to produce a translation as an insult to Jackson’s Korean, prompting everyone, including Jackson, to laugh. Jisoo follows up my comment with another evaluation of Google. Jisoo points out that Google is “very clever” (line 22), joking that Google recognizes Jackson as a “foreigner” (line 23), that is, not a Korean, and so translates his speech into Korean rather than English. Jisoo thus reframes the interaction once again, this time by attributing Google’s choice of the wrong language to its “clever” ability to recognize that Jackson is a foreigner (which is a fact) rather than his poor Korean (which is a negative evaluation). At the same time, her laugher (line 19) also positively acknowledges my attempt at humor, thereby allowing her to construct solidarity with both of us.
Excerpts 3 and 4 not only illustrate how evaluations of Google and its dialog are used to reframe out interactions with the device into opportunities for constructing solidarity and humor between family members, but also how often our evaluations associate it with intelligence and intent. This also applies to Excerpts 1 and 2, which feature similar instances of evaluation. The theme of intelligence is featured in Excerpt 2, where Jinyoung calls Google “stupid,” and Excerpt 4, where Jisoo does the opposite by evaluating Google as being “very clever.” These evaluations, while expressing different assessments of Google, both associate its ability to respond appropriately to users with its “intelligence.” Excerpts 1, 3, and 4 all feature attributions of intent to Google (and in doing so, arguably attribute it with social intelligence). In Excerpt 1, Jackson says that Google must be “profiling” me, while in Excerpt 3, Jinyoung accuses Google of “disqualifying” and “rejecting” him. I attribute intent to Google in Excerpt 4 when I reframe its mistaken translation of Jackson’s Korean utterance as an insult to his language skills.
The recurrence of such evaluations in my data highlights how commonly such traits are associated with AI and illustrates how our everyday practices in talking about such devices reflect big-D Discourses about AI as intelligent machines. As mentioned, Gordon (2013) notes that speakers draw on their knowledge schemas to interpret technological objects in any given conversational moment; in the data I have analyzed, it is clear that the association of Google with intelligence and intent is a part of my family’s knowledge schema about the device and the technology on which it operates. Of course, such knowledge schemas are not unique to my family; as previously discussed, Reeves and Porcheron (2023: 574) remark that AI systems are routinely agentified 2 in popular discourse, with claims of AI chatbots displaying “understanding” or “knowledge” of natural language being a key part of how they are marketed to the public. The way my family draws on such Discourses of AI is reminiscent of how Tovares (2012) has shown family members incorporate public discourse, such as references to TV shows, into private interaction Given how frequently the capabilities of conversational AI are exaggerated in relation to the notion of “AI as intelligent machines” in public discourse, then, it is perhaps unsurprising how often similar language is used by my family members in their evaluations of Google’s dialog, albeit most often in a joking manner. We do not attribute such qualities to Google because we believe that it is a sentient being participating in social interactions as humans do, but our everyday use of such language in association with Google nonetheless draws on and contributes to the pattern of attributing agency and intelligence to such devices.
Conclusion
To conclude, I have analyzed my family’s interactions with Google Assistant to examine how such interactions are integrated into our interactions with each other. Drawing on previous research on framing in family discourse (Gordon, 2008, 2013; Tannen, 2004, 2006), I have demonstrated how family members, through reframing and blending frames, connect their interactions with Google Assistant to interactions with each other. I have shown that evaluations of Google Assistant and its dialog are a frequent strategy used for the management of frames, and illustrated how such evaluations reframe interactions with the device into opportunities for constructing humor and solidarity among family members. I have also discussed the frequent attribution of intelligence and intent to Google Assistant in family members’ evaluations of its dialog. Specifically, I have argued that the use of such language to characterize conversational AI draws on the knowledge schema of “AI as intelligent machines,” and needs to be understood in the context of the social settings in which our interactions with conversational AI are situated.
My analysis of how themes of intelligence and intent come up in my family’s evaluations of Google sheds light on the multifaceted role that aspirational big-D Discourses about AI play in our everyday interactions with and about such devices. On the one hand, the practice of attributing intelligence to AI can have harmful effects in the sense that it misleads users on what the technology actually does that is behind products labeled as using “AI.” In a broader sense, one could also argue that it draws connections between the AI tools we use in everyday life and the idealized vision of AGI popularized through science fiction; this vision, of course, has new political meaning in the current era of AI hype, as AI evangelists push an ideologically-laden narrative of progress in which they envision the eventual emergence of a superintelligent AI that can solve all of the world’s problems, which, according to them, makes it worth overlooking the many problems with the AI industry (such as copyright issues, environmental impact, racial and linguistic bias, and so on) (Benjamin, 2024). Yet, my analysis suggests that such attributions of intelligence are not a straightforwardly problematic practice. Such evaluations serve as a key resource in everyday discourse by which my family members not only integrate our interactions with Google into our interactions with each other, but also use our interactions with Google as a means to build a sense of connection with each other. By bringing the concept of framing as a tool to better examine the layers of interaction involved in the use of conversational AI in group settings, my analysis clarifies the role that attributions of intelligence play in our interactions with and about such technologies.
The scope of this study is limited in that it draws on a small sample of one family’s interactions with one type of conversational AI device in two languages. Nonetheless, my analysis contributes to existing research by bringing together insights from the sociolinguistic study of family discourse and research in the field of human-computer interaction on the use of conversational AI in social settings. As mentioned, studies such as Porcheron et al. (2018) and Reeves and Porcheron (2023) have highlighted the need to consider how users’ interactions with AI unfold in relation to the social settings in which they take place. I build on their argument by drawing on research on framing in family discourse; my analysis shows that framing can be an especially helpful conceptual tool in discerning the complexity involved in AI interaction, as it provides an analytical framework by which we can pinpoint what the functions of paralinguistic and linguistic strategies are in relation to specific interactive frames. I also extend their discussion of the social constituted nature of AI interactions by demonstrating how processes of meaning-making in such interactions are shaped not only by the immediate social context, but also by big-D Discourses about AI, in this case of “AI as intelligent machines.” Finally, I add to research on the use of technology in face-to-face interaction, notably İkizoğlu’s (2019) study on users’ interactions with a voice translation app, by considering users’ interactions with and about such technologies as distinct yet interrelated layers of interaction.
Important steps have been taken in recent years in bringing together the discourse analytic study of conversation with research on conversational technologies, such as in the special edition of Discourse & Communication by Stokoe et al. (2024) on this very topic. I argue that in further exploring and expanding this intersection, we should consider how discourse analysis can not only be leveraged to improve such technologies, but also shed light on the ways in which they are being talked about and understood in everyday life, so as to establish a more holistic understanding of users’ interactions with conversational AI. By drawing on previous research on framing, this study illustrates how the link between AI and agentification is reinforced in everyday talk, by demonstrating how even nonsensical or erroneous linguistic output produced by conversational AI can be made meaningful by human speakers. More broadly, research in this vein can help us understand phenomena surrounding AI such as “agentification”, i.e., the process by which “AI systems’ supporting societal and social infrastructures are routinely erased from the discussion” (Reeves & Porcheron 2023: 574) in the service of promoting its capabilities as an “autonomous” technology. My analysis sheds light on a related process, that is, the attribution of intent to AI for meanings that are in fact produced through the surrounding commentary of human speakers. To users of conversational AI, the findings of this study— and of future work along these lines— may provide an opportunity to critically reflect on claims about the “human-like” capabilities of AI and the work that users do that makes AI-generated output seem more impressive. To technology developers, this insight could be of use in informing their understanding of how users develop social meanings about AI. For example, applying the analytical lens used in this study to online forums where users talk about their experiences with AI chatbots (e.g., Reddit’s r/MyBoyfriendIsAI), could shed light on the process by which users come to rely on chatbots as a meaningful source of emotional support. Such insights may help contextualize broader concerns raised in public discourse about overreliance on AI and its potential social and psychological consequences.
Footnotes
Appendix: Transcription conventions
Source: Transcription conventions adapted from Tannen, Deborah, Shari Kendall, and Cynthia Gordon (eds) (2007) Family Talk: Discourse and Identity in Four American Families. New York: Oxford University Press.
Acknowledgements
I would like to thank Cynthia Gordon, Deborah Tannen, and Stuart Reeves for their invaluable guidance and feedback, as well as the editors and anonymous reviewers for their thoughtful suggestions. All remaining errors are my own. This study was conducted as a part of my doctoral dissertation at Georgetown University, and a previous version was presented at the 2024 National Communication Association Convention.
Ethical considerations
Ethical approval for this study was obtained from Georgetown University’s IRB (Protocol ID: STUDY00005630).
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: The author declares no commercial or financial conflicts of interest. However, the author acknowledges that the participants were members of the researcher’s family. This was disclosed and approved as a part of the ethical review process.
Data availability statement
The data used in this study are not publicly available due to privacy restrictions involving human participants.
