Abstract
This study explores the efficacy of keystroke logging as a method to qualitatively investigate the synchronous processes of discursive interaction through mobile devices as individuals go about their everyday lives. Heeding cautions from Boase (2013) concerning software variability across mobile technologies, as well as challenges from Ørmen and Thorhauge (2015) to use log data for qualitative research, our study offers one such methodological roadmap for observing—from the software side—the complex entanglements of humans and mobile technologies as they engage in mediated discourse. Our study draws upon keystroke analysis from the tradition of writing process research (Leijten & van Waes, 2013; Wengelin, 2006), as well as posthumanist methodologies for observing cybernetic interactions (Giddings, 2014), and extends Farman’s (2012) argument that asynchronous forms of mobile communication, such as text messaging, are performatively synchronous, since interlocutors are pulled toward embodied copresence via the mediated space. In doing so, we present a preliminary study of methods for directly observing how discursive processes manifest in-the-moment as a complex flow between human, machinic, and spatial components in a network assemblage.
Keywords
Mobile communication research has advanced understandings of mediated discursive practices, particularly the ways in which ubiquitous mobile technologies alter how individuals interact with social networks (Humphreys, 2010, 2012), experience space and locationality (de Souza e Silva, 2013; Farman, 2012), and perform and maintain relationships with other individuals (Cui, 2015; Hall, Baym, & Miltner, 2014; Ling, 2008). While scholars have qualitatively demonstrated a myriad of discursive practices that individuals engage in continuously throughout their everyday lives (Mihailidis, 2014), discursive understandings are methodologically limited to observing asynchronous products picked from ongoing streams of mobile interactions. Scholars have developed a robust understanding of how certain text-based content discursively functions (e.g., Baron & Ling, 2011; Tagg, 2012; Thurlow, 2003); however, we have little understanding of how this content is produced in-the-moment and therefore reflects individuals’ everyday embodied and mediated experiences through mobile technologies.
This study explores the efficacy of keystroke logging as a method to qualitatively investigate these in-the-moment acts—what we call synchronous processes—that are involved in producing discursive interaction through mobile devices. Heeding Boase’s (2013) concerns regarding software variability across mobile technologies, as well as recognizing the challenges Ørmen and Thorhauge (2015) identify for using log data for qualitative research, our study offers a methodological roadmap for observing, with software we designed, the complex entanglements of humans and mobile technologies as they engage in mediated discourse. Our methodology draws upon keystroke analysis from the tradition of writing process research (Leijten & van Waes, 2013; Wengelin, 2006), and is informed by posthumanist methodologies for observing cybernetic interactions (Giddings, 2014). We extend Farman’s (2012) framework for a performative understanding of mobile communication, particularly the notion that individuals who are in the act of producing asynchronous messages, such as text messaging, are engaging in synchronous performances that directly reflect the discursive context. This study therefore explores methods for directly observing how discursive processes manifest in-the-moment as a complex flow between human, machinic, and spatial components in a network assemblage.
We situate our study amongst ongoing research concerning mediated discursive practices in mobile communication technologies as performative, and discuss our methodological inspirations from keystroke logging analysis and microethology. Next, we present the design of our present study, with combined results and analyses. We follow with a brief discussion of the implications of our study’s findings and the wider application of our methods. In doing so, we hope to push mobile communication research towards a more nuanced understanding of the spaces, temporalities, and bodies engaged in everyday smartphone use.
Background
Our study begins with a reconsideration of the concepts synchronous and asynchronous, particularly as they refer to discourse. In this framework, synchronous refers to discourse that is transmitted through a modality that is unique to specific moments in time (e.g., speech), and asynchronous refers to discourse that may be transmitted through a modality that may be repeatedly accessed (e.g., a written record). Written texts are traditionally and temporally considered to be asynchronous media that reflect the time and space in which they were composed (Stiegler, 2008). Since many forms of text-based mobile communication involve discrete information exchange (e.g., text messaging, instant messaging (IM), Snapchat), individuals receive a finalized message, but know little about how their interlocutors composed said message. In other words, while the receivers of a text-based message encounter an asynchronous product, the sender of the message engaged in a synchronous process to compose that message.
We have therefore adopted Jason Farman’s (2012) argument that asynchronous media technologies downplay or omit the embodied performance of mediated interaction, or the space in which interlocutors reaffirm copresence (Thurlow, Lengel, & Tomic, 2008). In this framework, Farman argues “embodiment is always a spatial practice,” and that mobile media extends the sensory and spatial reach of an individual, what he calls the “sensory-inscribed” body (2012, p. 19). This means that interlocutors may simultaneously negotiate far-off social spaces while communicating via mobile devices (Humphreys, 2010), and that they remain committed to maintaining copresence despite spatial and temporal distances (de Souza e Silva, 2013; Ling, 2008; Ling & Yttri, 2001). Farman (2012) therefore argues that even if the media is asynchronous, mediated interaction is performatively synchronous. According to Karen Barad (2007), this notion of performative means that living is an active ongoing process that occurs amidst many bodies, beings, and things—what may be called an assemblage. Barad argues that for scholars to adopt a performative approach, they must “focus inquiry on the practices or performances of representing, as well as on the productive effects of those practices and the conditions for their efficacy” (Barad, 2007, p. 28). Farman’s and Barad’s approaches to performativity therefore require attention to what bodies—human and machinic—do in-the-moment while engaging in discursive practices. How individuals synchronously maintain an ongoing negotiation of embodied spaces, rituals, and discursive practices as they compose as well as await message content on mobile devices is therefore critical to the performative approach.
Even though many types of embodied performances may be observed through ethnography (see Baym, Zhang, & Lin, 2004; Humphreys, 2010; Ling, 2008), it bears asking: How might written discourse be observed from a synchronous perspective? Doing so might not only shed new light on the embodied performances involved in writing and sending text messages, but also provide a more holistic picture of how time is made to matter for asynchronous media.
In the field of writing process research, the distinction between asynchronous and synchronous issues encompasses the point of observation. Specifically, asynchronous data are collected from observations after writing has concluded, for example, after an individual sends an email; synchronous data are collected from observations of writing as it occurs in real time, for example, while an individual is composing an email (Janssen, van Waes, & Van den Bergh, 1996). While both types of data are crucial for fully understanding how individuals meaningfully produce texts, synchronous data provide a richly detailed understanding of the psychomotor processes through which individuals negotiate social and physical forces while translating and transcribing thought into text (Hayes, 2012; Leijten & van Waes, 2013).
We suggest that incorporating this understanding of synchronous observations of writing requires a focus on the processes through which individuals compose texts on mobile devices. Ekman (2015) argues that various forms of mobile communication, like Snapchat, obfuscate received messages so they are less artifact and more an ephemeral embodiment of copresence. This suggests that asynchronous observations may be overvalued, and mobile communication research should also focus on how individuals work to perform copresence in between the sending and receiving of messages. How individuals use virtual keyboards to translate and transcribe discourse through mobile devices may be indicative of individuals performing various discursive practices, such as reply norms (Laursen, 2005), hyper-coordination (Ling & Yttri, 2001), relational maintenance (Spilioti, 2011), and communicative intimacy (Thurlow, 2003).
Synchronous observation, however, poses a daunting challenge for mobile researchers. While text-based messaging on computers, such as instant messaging, has afforded stable vantage points for in-the-moment observation (see Lewis & Fabos, 2005), the mobile in mobile messaging means that direct and synchronous observations require a laborious level of agility and time commitment, in addition to potentially impeding participants’ mobile activity. Boase (2013) suggested that a mobile device’s software may provide useful evidence of communicative practices that individuals engage in across devices, and Ørmen and Thorhauge (2015) similarly encourage the use of log-file data in tandem with ethnographic methods. These suggestions are reminiscent of Giddings’ (2014) microethology, which calls for attention to human and machinic agencies while observing human interaction with digital media. The inclusion of log-file data is akin to treating the mobile devices as participants in qualitative research.
Our study therefore posits that collecting keystroke log-files—a chronological record of each individual press of a key on a keyboard—from the virtual keyboards of mobile devices will provide synchronous observations of mobile discursive practices. We draw upon keystroke analysis from writing process research (see Leijten & van Waes, 2013) to focus on examining writing processes via the time elapsed between input actions (pause analysis), as well as the timing and contexts in which individuals delete textual content that they had previously inputted (revision analysis). While pause and revision analysis typically examine how individuals compose longer texts (e.g., workplace written reports), these analyses are theoretically indicative of how individuals actively and consciously articulate communicative content through text (Leijten, Van Waes, Schriver, & Hayes, 2014). For example, Van Waes, Leijten, and Quinlan (2010) demonstrated that competing cognitive resources are evidenced by pause-based disruptions during composition, showing that the act of deleting and revising composed text signals a complete cognitive shift to one of the competing tasks (i.e., carefully rereading text to search for misspellings). Leijten et al. (2014) further argue that these attentional shifts may occur due to the competing psychomotor, social, physical, and technological demands of the writing task. Based on these findings, the sociophysical space, response times, and even familiarity with the writing interface may all present during pause-based disruption or increased textual revision. In the following section, we detail our study’s methods for exploring synchronous mobile discursive practices through keystroke logging analysis.
Methodology
In order to guide the highly experimental nature of our inquiry, our study was primarily informed by the question, How does keystroke logging inform observations of synchronous mobile discursive practices? We sought to develop a first-of-its-kind mobile keystroke logging application as well as an experimental design to simulate real-world mobile discursive interaction through a smartphone device.
Drawing upon recent innovations in keystroke analysis from writing process research (see Leijten & van Waes, 2013; Leijten et al., 2014; van Waes et al., 2010), we developed a custom chat application for use on an HTC Incredible One running Android 4.0 mobile operating system (see Figure 1). This chat application, which functioned as a one-to-one messenger service between the phone and the application’s servers, simulated text-based mobile messaging (e.g., text messaging or IM) between participants equipped with the phone at one end and the researchers at the other end (see Figures 2 and 3). The application simultaneously logged the synchronous keystroke data as well as asynchronous chat transcripts (i.e., message content that was sent between phone and server).

An image of the HTC Incredible One running the custom chat application.

An image of the application design.

An image of the chat application from the server side.
Our experimental design attempted to simulate real-world mobile discursive practices; namely, the process of walking through a sociophysical space while simultaneously texting on a mobile device. Mobile discourse scholars have noted that the time spent replying to text messages from certain network ties (Laursen, 2005; Ling & Yttri, 2001), as well as the selection of marked textual content (Baron & Ling, 2011; Spilioti, 2011), signal the complex processes individuals engage in as they negotiate the local, distant, and virtual spaces they are networked to while using their mobile phone (Humphreys, 2012). To qualitatively understand how individuals coordinate simultaneous communicative tasks in overlaid social spaces, our study’s use of keystroke analysis will demonstrate how participants manage reply norms (Laursen, 2005) and intimacy orientations (Thurlow, 2003). In the following sections, we detail our study’s procedures as well as our methods for analysis.
Procedures
Considering the highly experimental nature of our project, we piloted our methodology by conducting a set of qualitative case studies (N = 5) that were approved by the Institutional Review Board at the mid-Atlantic university at which this study was conducted. Participants were recruited via snowball sampling methods through formal and informal social networks, and they were additionally encouraged to bring a friend or colleague for participation. Each case study was conducted at the main university library that is largely used by university students, but is open to the public. Once participants consented, they were provided with the HTC smartphone and a GoPro HERO video camera strapped to a head-mount. The smartphone was preprogrammed to communicate via text chat directly with the researchers, and included a keystroke logging application to capture participants’ keystrokes when they typed with the phone’s QWERTY-style virtual keyboard in the chat application (see Figures 1–3). Each participant was then provided a set of tasks to complete within the library that we communicated to them via the smartphone. Before any tasks were initiated, participants were asked a series of interview questions, and after they completed their tasks they were asked another series of follow-up questions.
Our study design consists of three tasks, which have a series of communicative goals associated with them. In Task 1, participants were instructed to walk from a café space on the main floor of the library to the fifth floor via the elevators and to text the researchers once they arrived. Participants were then asked to search for a book in the stacks on the fifth floor. Following that, they were asked to text the researchers once they found the book. Participants were then instructed to relocate to the first floor of the library and go to the “Learning Commons,” a crowded multipurpose space that is frequently used for studying, playing videogames, meetings, and more. Once there, participants were instructed to find a dry-erase board and draw a picture. When each participant completed the final task, they were asked to return to the main floor café.
Task 2 operated in parallel with Task 1. Participants were asked questions in a semistructured conversation with the researchers via the text-chat application. Our conversations were designed to parallel Task 1 requirements, entangling participants in a multidimensional text dialogue that balanced multiple intimacy orientations (Thurlow, 2003). These conversations involved questions that centered around the texting habits and attitudes of texting of the participants, such as: “How often do you use text messaging or IM chat applications on your phone?”; “Who do you text with typically?”; “What do you like about texting?”; and, “If you receive any texts during this study, feel free to answer (just remember that the camera will be rolling)”. Participants were asked about their habits surrounding smartphone usage, such as: “Do you own a smartphone?”; “What do you use your smartphone for?”; and, “How long did it take you to get used to using a smartphone compared to another type of phone?” The researchers asked participants about their academic writing habits, such as: “How frequently do you write for your academic work?”; “Do you have a tablet that you use for schoolwork?”; “What sort of schoolwork do you do on your tablet?”; and, “Do you ever try to write on your tablet? Do you like that?”
Coding and analysis
The audiovisual data recorded from the GoPro was coded using ELAN Version 4.8.1 (Lausberg & Sloetjes, 2009). While the social environment and procedures allow for variable types of interactions to occur, the primary focus of the audiovisual data was the coordination between mobile movement through the social environment and the participant’s focus on textual production (see Figure 4). Our coding coordinated participants’ attention being drawn to as well as drawn away from the mobile device, the texting discourse with the researchers, performance of the three tasks, and any other level of engagement with the social space of the library.

A screenshot of coding conducted in ELAN with video data.
Using ELAN’s Version 4.8.1 interface for annotating synchronous data along multiple simultaneous dimensions, the audiovisual data were coded with a mutually exclusive and exhaustive scheme to track how participants engaged with the smartphone device (the PHONE dimension), how participants physically moved throughout the library (the MOTION dimension), and what the potential social surroundings were as indicated by the camera footage (the SOCIAL dimension). In the PHONE dimension, we coded for when participants composed text messages, navigated the smartphone user interface, gazed at the smartphone device with the screen on, gazed at the smartphone with the screen off, placed the phone down or in a pocket, or used a completely different application on the phone. In the MOTION dimension, we coded for when participants were walking, climbing stairs, riding the elevator, pacing or walking without an apparent direction, kneeling or sitting, or drawing. In the SOCIAL dimension, we coded for when participants were in a physical/social position that could have allowed for social discourse.
The primary discursive forms analyzed for this study focus on temporal factors, particularly the elapsed time (in milliseconds) between the initial press of a finger on a discrete key on the QWERTY virtual keyboard (i.e., a keystroke), on whether the keystrokes corresponded to alphanumeric textual inputs, on the deletion of inputted text, and on the use of suggestive text. The applications we developed (e.g., the Android app and Windows messaging server) communicated using XML-like markup to organize the information in each message, and the server performed two tasks when a session was saved: first, the messaging packets were aggregated in a log file (Table 1); and second, the keystroke data were parsed from each message and saved in individual comma-separated value files (CSVs; Table 2). CSV files can be read by Microsoft Excel and are self-documenting, like XML files, because they are human readable. 1
A message packet from a session.
Parsed keystroke data from a collection session seen in Table 1.
We then produced reports from the data we captured by analyzing the keystroke values and comparing what was written to what was finally sent (e.g., the “MESSAGE” value vs. the encapsulated “KEYS” values). As Table 3 shows, we assigned the first keystroke of a session a “0” duration, as that value had no logical antecedent. Each subsequent keystroke’s pause value was measured by determining the time between its value and its antecedent’s value. In this way, each spreadsheet noted the time (in milliseconds) at which each key was pressed (Time ms), the specific key (Event) that was pressed, whether it was determined to be an alphanumeric or function key (Event type), the time that elapsed between the previous keystroke (Pause ms), and the time elapsed for the entire session (Duration ms).
Compared to long-established keystroke analysis tools for computer-based writing such as Inputlog (Leijten & Van Waes, 2013), our system for collecting and analyzing keystrokes from the virtual keyboard was limited to logging the time at which a key is pressed, and what key is pressed. While a pause is commonly defined as an interkey interval between the release of one keystroke and the initial press of the next keystroke (Leijten & Van Waes, 2013), for the purposes of our study, a pause must be defined by the intervals of time that elapsed between the initial press of subsequent keys when typing with the virtual keyboard. For example, as seen in Table 3, since “d” was pressed at 15,681,695 ms and “k” was pressed at 15,682,104 ms, the interkey interval is therefore calculated by subtracting the key-press for “k” from the key-press for “d,” which is 409 ms. This limitation in our definition of a pause was necessary in part because of mechanical distinctions between a virtual keyboard and a traditional computer keyboard; however, as we will argue in the Discussion section, future directions of our research will account for this distinction.
All graphical visualizations and analyses were conducted in RStudio v0.98b. These data contextualize the composition process of texting discourse, consisting of transcripts of discourse between the researchers and participants, and audiovisual recordings. The pre- and postinterviews provided qualitative data to help inform each individual case study in terms of the participants’ “typical” texting and mediated mobility habits.
Results
The five observational case studies were conducted in April 2015 at the university’s main library. All participants were female graduate students at various stages of progress and from various programs. Three of the five participants used Android devices, and two used Apple devices. Four of the five participants reported having used a smartphone as their primary phone for “several years.” Participant K was a “late adopter” of smartphones and had only begun using her smartphone 3 weeks prior to her participation in our study.
The time period in which the sessions took place happened to correspond with finals week of our university’s spring semester, and there was a higher level of activity in the library than had been previously considered. The text-chat application (which relied on the university’s WiFi network) had occasional issues maintaining a stable connection to the local networks, which briefly interrupted the observation study involving Participant A (she returned to the café to check in with the researchers before resuming). A fire alarm required evacuation of the building, interrupting Participant O’s session, causing the loss of video footage for Participant B. Only three participants completed all tasks in the study without interruption (Participants B, C, and K); video footage is available for four of the participants (Participants A, C, K, and O); keystroke logging data are available for all participants. The results of the keystroke logging data will be discussed first, followed by a discussion of the video data, and then a discussion of insights gained from coordination of the multimodal data.
Keystroke logging on a smartphone
Examining the keystroke logging data obtained from participants in this study revealed both technical achievements and limitations regarding the level of granularity our methods captured. Specifically, the keystroke logging data successfully captured the synchronous composition of textual content, as well as the recursive and discursive nature of that composition. Across all participants, the keystroke logging data offered a synchronous visualization of when participants texted, by how much, and for how long they texted. Figures 5, 6, and 7 demonstrate the input of alphanumeric characters over the time of the study, and in each graph text input tends to cluster together, representing discrete periods of time when each participant was composing text. The patterns of these clusters may correspond with specific types of interaction. In Figure 7, early clusters appear much shorter in length, representing texts intended to correspond to procedures of Task 1; however, the long cluster in the middle of the graph corresponds to Task 2 wherein the researchers asked if the participant uses her phone to write for professional or academic purposes. Figure 5 similarly exhibits longer clusters of texting, which happens to correspond with C’s long ride in an elevator. Participants O, A, and C thus illustrated a curvilinear pattern in their composition, suggesting that their compositional habits adjusted due to the circumstances of the study. In contrast, Participants B and K exhibited a straightforward linear pattern of composition (particularly K, see Figure 6), suggesting a steadier interaction between texting and sociophysical environments.

A timeline of Participant C’s keystroke input.

A timeline of Participant K’s keystroke input.

A timeline of Participant O’s keystroke input.
Examining the types of keystroke events and the time between them affords a picture of revision patterns in texting. In Figure 8, time between alphanumeric keystroke events, deletion of text (marked “backspace”), and use of suggestive text (marked “auto-complete”) demonstrates the deletion of existing text largely occurs at variable points within the composition of a single message between 150 and 300 milliseconds, and that they might occur in tandem with reading processes as texters make revision decisions. These processes are demonstrated in Tables 4 and 5, showing a brief selection of the final transcript of the dialogue with Participant O as well as the keystroke-by-keystroke construction of portions of the transcript.

A subsection of Participant O’s keystroke input, with pause duration (ms).
Sample transcript with Participant O.
Note. Boldface added for emphasis.
Visualization of the keystroke composition of a single clause of text from Participant O.
In Table 4, we can see how Participant O clarified and then took the time to compose a longer formal answer to a Task 2 question (this was the long cluster of keystrokes seen in Figure 7). She took over 2 minutes to compose the message to answer the Task 2 question “How frequently do you write for your academic work?” as seen in Row 5; however, within approximately 14 seconds she answered a Task 1 question that was sent while she was composing her Task 2 answer in Row 6. Her message in Row 5 contains the only instances of personal subject pronouns, which are omitted in Row 6. This reflects Thurlow’s (2003) texting maxims that compel texters to be syntactically brief and quick, depending upon the relational and social orientation of the discourse. Figure 8 visualizes the construction of the text sent in Row 5 from Table 1, and shows that while deletion of text occurred throughout the construction of the entire message, a large degree of revision occurred early in the message (see Table 5), thus demonstrating O’s attempt to manage the communicative orientation imposed by the question.
Revision analysis was conducted post hoc by transposing keystrokes into deletion segments based on when participants pressed BACKSPACE (e.g., the segment “maybe⇤⇤⇤⇤⇤notsure” typed by Participant B). In total, of 66 deletion segments that were observed from the keystroke data, 60 of them were intraword deletions (i.e., to delete a specific character or delete an entire word) with a median pause after the deletion of 669 ms, and six deletions were interword deletions (i.e., to delete multiple words) with a median pause after the deletion of 972.5 ms. Furthermore, the deletion segments predominantly concerned Task 2 responses (48 segments) as opposed to Task 1 responses (18 segments), and the median pause after intraword deletions were 688 ms for Task 2 and 480 ms for Task 1. Participants were therefore more likely to delete and revise textual content related to Task 2, and such deletions were typically intraword deletions with longer pauses (see Table 6).
Deletion segments.
Visual data of texting
After the audiovisual data were annotated in ELAN Version 4.8.1, the synchronous annotations were visualized in RStudio to construct timelines of events. The timelines for Participants C and K (Figures 9 and 10) demonstrate that participants differed in terms of how they balanced the sociophysical space, the tasks, and the texting discourse. As visualized in Figure 9, Participant C texted almost continuously while waiting for and riding in the elevator, regardless of the presence of others. In fact, she reported that she was so absorbed in texting while in the elevator that she missed the 5th floor. Further, as visualized in Figure 10, Participant K appeared to avoid texting while she was physically moving around the library, in contrast to C, whether she was walking or taking the elevator. At one moment, while walking briskly, Participant K received a text from the researchers and immediately stopped and stood still in the middle of the “Learning Commons.” She then appeared to read the message from the researchers, composed her response, and then resumed walking.

A timeline of Participant C’s session as coded by audiovisual data.

A timeline of Participant K’s session as coded by audiovisual data.
Situational embodiment and texting
Each participant exhibited multiple moments where local and distant demands were mediated through embodied interaction with the smartphone device. For example, when Participant C was waiting for and then riding in the elevator, she was texting almost continuously, largely when other individuals were present in the elevator as well. During this time, she balanced coordinating the errands involved in Task 1 with the personal questions involved in Task 2. She texted the most to the researchers, both in terms of time spent composing and in keystrokes inputted (see Figure 5). In contrast to when she was navigating through the fifth-floor stacks and the “Learning Commons,” C was much briefer with her texts for both Tasks 1 and 2.
Such moments are similarly made apparent for Participants K and O. When Participant K stopped in the middle of the Learning Commons, she balanced both Tasks 1 and 2; the message that stopped her in her tracks continued a Task 2 conversation about her being used to her new smartphone, and then provided further instructions for Task 1. K’s response engaged in Tasks 1 and 2 within the same sentence: “still getting used to it and here.” Participant O, as already discussed, navigated both Tasks 1 and 2 after she arrived at the fifth-floor stacks. In contrast to Participants C and K, Participant O was very elaborate with her response to a Task 2 question but very brief with her response to a Task 1 question. That she texted from between two bookshelves—and mentioned during her follow-up interview that she felt awkward in that moment— demonstrated the effects of the discomfort our apparatus caused to the participants on the local and distant demands we asked them to achieve during our study.
While not directly engaging the tasks of our study, Participant O’s reaction to the fire alarm being pulled during her session had consequences for our discussion of situational embodiment. According to Farman (2012), a “live event is completely altered by . . . mediatization” (p. 97). After the fire alarm was pulled, Participant O continued texting and looking at her phone while moving toward and navigating the building’s stairwells. GoPro footage shows that she continued looking to the phone for prompts (time marker 0.01:54), and messaged us for information. While the last message we received from her was still task-oriented, we had, by the time of the alarm, stopped the server and ended the recording session, leaving the library in the process. The GoPro footage revealed that O continued to text us as she evacuated the library: she made her way down the stairs and continued to look for us through the application (time marker 0.01:30), and then checked her own mobile device (0.01:40) for information. O’s engagement with the phone while evacuating the library, like C missing her floor in the elevator, demonstrates that interactions with us over the mobile device had “become the site of the primary engagement” (Farman, 2012, p. 100).
Discussion
The guiding question of our methodology was: How does keystroke logging inform observations of synchronous mobile discursive practices? In combining ethnographic methods with mobile keystroke logging methods, we witnessed our participants’ embodied performances (Farman, 2012) as they negotiated the sociophysical space of the library, the discourse with the researchers, and the cybernetic interface of the mobile device. Synchronous observation of these performances provided rich understandings of how individuals maintained copresence with near and distant social actors through mobile communication. Specifically, our study yields two crucial observations: that communicative orientations may involve distinctive in-the-moment discursive practices; and that the embodied processes of writing and physically moving about a social space are intimately bound.
As participants responded to text messages in Task 1, which concerned instructions for running errands within the library, or Task 2, comprising biographical questions, keystroke analysis revealed distinctive timing patterns for writing. Participants were faster at composing and responding to Task 1 messages, which included less textual content; whereas, for Task 2, participants often took more time to respond, included more textual content in their messages, and were more likely to revise textual content by deleting and rewriting text. As Participant O demonstrated (see Tables 4 and 5), switching between Task 1 and 2 communicative orientations is signaled by greater deletion and revision activity, and increased pauses between keystrokes. In other words, the keystroke data evidenced a qualitatively distinct level of attention to textual content for more intimate orientations. While the brevity of the messages and short amount of time between responses have been previously documented as discursive cues (see Laursen, 2005; Thurlow, 2003), that the amount and timing of keystroke input activity distinguishes communicative orientations is a new insight.
The coordination of keystroke data and audiovisual data—akin to human and machinic vantage points—provided us with detailed evidence that mobile discursive practices and embodiment are intimately entangled with one another. Observations demonstrate that each participant simultaneously negotiated the sociophysical space of the library as well as the discursive space demanded by the receiving and sending of text messages, and each participant managed these demands in distinctive ways. Whether participants mishandled an elevator ride (Participant C), hid from social view (Participant O), or stopped in their tracks in a social space (Participant K), participants engaged in different strategies for managing their simultaneous presence in the library and in discourse, and this varied with each individual and discursive task. Because complex keystroke activity is evidenced by use of the backspace key and pause patterns for Task 2 messages, which was not the case for Task 1 responses, we conclude that more attention to word choice and spelling indicates participants were using the sociophysical space to manage their attention within the immediate surroundings. In other words, our observations from the combination of synchronous data channels evidence how varying discursive processes for maintaining copresence flow through individuals’ embodied performances.
As previously discussed, Farman (2012) has argued that embodiment is a “spatial practice” that is evidenced in mobile discourse through the “sensory-inscribed” body (p. 19). This notion of the “sensory-inscribed” body unites “the sensory interaction between the limits of my body, the space it contextualizes it, and the interaction with your body” (Farman, 2012, p. 26). In other words, using mobile devices for discursive practices alters or extends the senses that are inscribed into the embodied performance, which Farman suggests is casually observable when individuals filter in or filter out certain sensory information over others while using mobile devices. For example, responding to a text message may divert attention away from noticing which floor an elevator has reached. Based on our observations of how our participants synchronously composed text messages, we therefore suggest that the “sensory-inscribed” body may additionally account for the nuances and complexities of discursive practices and maintaining copresence. Further study may lend credence to these observations. In the following section, we discuss our study’s limitations as well as suggestions for future directions.
Limitations and future directions
Based on our preliminary study, we have made four observations that speak to both the limitations of our study as well as suggestions for future research. Specifically, these points speak to the use of the GoPro camera, the complementarity of our multichannel data sources, our method of keystroke logging, and the “naturalistic” conditions of our study.
First, we recognize that our requirement for participants to wear the GoPro configured their social contexts by increasing their self-conscious feelings, as participants noted in their poststudy interviews. The GoPro likely dissuaded participants from remaining in social spaces long, which can explain why Participant A took the stairs rather than the elevator, why Participant O stood between bookshelves, and why Participant C might have missed the fifth floor while riding the elevator. Maintaining focus on the mobile device therefore may have become a way for participants to avoid the gaze of others. While future study could certainly involve less intrusive methods for embodied observation, we nevertheless argue that participants incorporated the GoPro into their embodied performances, as they did with the sociophysical space of the library. Nevertheless, triangulating the audiovisual data, keystroke data, and interview data allowed us to contextualize individual participant experiences, that is, to understand why C missed the fifth floor, and why O stood between the bookshelves. For future study that involves observation over longer periods of time, additional interviews would better contextualize synchronous data, and perhaps require less use of video data.
Second, while we argue that our incorporation of multiple channels of synchronous data (keystroke logging as well as video recordings from the GoPro) yielded richly detailed observations, these channels were not entirely complementary. Each data channel originated from different devices, which therefore required the researchers to infer the points of calibration. While we are confident in our approximations, this is nonetheless based on an inference. Future research could resolve this by having both data channels come from the same device and calibrated to collect data based on input into the device. For example, any time a key is pressed, a screenshot or picture from the front-facing camera could be simultaneously collected. Such data would certainly result in video data from a different vantage point and of a different quantity, but would nonetheless be complimentary to the keystroke data.
Third, as discussed in our Methodology section, our applications were designed to be exploratory, and we were limited by the hardware and software constraints of the HTC Incredible One, which was running Android 4.0. As a result, while we used established keystroke logging tools, such as Inputlog, and recommendations for XML-structured logging tools (van Waes, Leijten, van Horenbeeck, & Pauwaert, 2012) as general guidelines, we did not design an application that was entirely within the confines of standardized practices. Since we are not aware of other keystroke logging applications intended purely for mobile smartphones, and we were nonetheless able to successfully log and analyze the initial presses of individual keystrokes, we argue that our study has made an important contribution to the study of mobile communication as well as keystroke analysis. In future study, we will therefore be prepared to design keystroke logging tools for updated mobile devices (i.e., Android 8.0), and to match recommendations from Van Waes et al. (2012).
Lastly, future study would involve not only a more robust participant pool, but an expansive and “naturalistic” means of data collection. We argue that the simulated conditions of our study demonstrated the value of keystroke logging for studying mobile discursive practices; however, we by no means claim our study is a “naturalistic” observation. We attempted to isolate and replicate certain everyday discursive conditions, but cannot account for all of the improvisational and creative practices that occur in individuals’ daily lives. A qualitative study that observes individuals over longer periods of time as they communicate with members of their interpersonal networks is an ideal context for incorporating keystroke logging methods.
Conclusion
Our study demonstrates experimental methods for observing synchronous mobile discursive practices, providing a detailed picture of how individuals use mobile devices as part of everyday mobility. Discourse, cybernetic interfacing, and sociophysical mobility were all intertwined as our participants navigated the library space to satisfy requirements that simulated everyday interaction. Future research that incorporates keystroke logging on mobile devices may well come to richer and more nuanced understanding of written communication as part of mobile discourse.
Our use of synchronous observation methods reveals an understanding of how mobile discourse interacts with sociophysical factors, allowing such discourse to be understood in terms of an embodied performance (Farman, 2012). Participants’ use of the library spaces and mobile device to maintain an ongoing copresence with the researchers involved an attentional balancing act between and amongst the overlaid spaces; like a metaspace (Humphreys, 2012), the mobile device’s software created log-file data which, when coordinated with more traditional ethnographic methods in the form of a microethology (Giddings, 2014), demonstrated that the embodied performance of texting is a synchronous practice.
Naomi Baron (2013) has argued that mobile devices challenge our understanding of written communication and discursive processes. She notes that incorporation of the QWERTY keyboard into touchscreen devices—which conjoins input and output components upon the screen—may put a strain on writing and reading processes, which may partially explain increased attention to automation, speech-to-text, and discursive norms (i.e., brevity) that reduce effort in writing. Baron therefore implies that mobile discourse is a shift in written communication. Leroi-Gourhan (1993) argued that the cultural development of writing would occur in three stages: the first stage involves “direct motive action of the hand with the hand tool” (i.e., handwriting); the second is where the hand supplies “indirect mobility” in which force transfers from the hand to mechanical processes (i.e., typewriting); and the third is where “the hand is used to set off a programmed process in automatic machines” (p. 242). This third stage is emblematic of writing on a touchscreen device, where the virtual keyboard is a software-based echo of writing on a mechanical typewriter. As our study evidences, suggestive and auto-text functions are intimately interwoven with the composition of nearly every message, representing a significant departure from our understanding of the agencies at play in mobile discourse.
Just as hands clasped pencils and mobilized graphite to paper, hands grasp smartphones and negotiate user interfaces, and the situation in which the author composes is shaped by their situational, sociophysical context. Text messaging describes a specific means of discourse, of the technology used, of the way the interlocutors are brought forth within an environmental context, and of the ways in which the interlocutors’ writing is revealed and shaped by the technology and context from which they write.
Footnotes
Acknowledgements
We would like to acknowledge Dr. Melissa Johnson for her guidance in the design of our study, as well as Dr. Jason Swarts for his insights into our analytical methods and rigor.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
