Abstract

Composition and Big Data, edited by Amanda Licastro and Benjamin Miller, offers the discipline a collection of 16 chapters that explain and discuss big data methods in composition studies. The editors use the book to advocate for “work that combines qualitative and quantitative methods, recognizing that data doesn't speak for itself, but must be spoken into and from, based on deep disciplinary knowledge” (p. 8). They attempt to not only broaden but deepen disciplinary methodological knowledge and provide carefully situated critiques of big data methods. This approach allows Licastro and Miller to emphasize exciting technological tools and how to use those tools while they productively problematize what these tools do.
The book is divided into four sections: “Data in Students’ Hands,” “Data Across Contexts,” “Data and the Discipline,” and “Dealing With Data's Complications.” Although each chapter follows the theme in its section, the book's pagination does not indicate the separation between sections, which can make thematic similarities harder to track.
Section 1, “Data in Students' Hands,” is the smallest, containing only three chapters. In Chapter 1, Trevor Hoag and Nicole Emmelhainz describe teaching undergraduate students to use “machinic collaboration” (p. 25) to assist in metacognitive textual analysis. They showcase an assignment asking students to use distant reading to make new interpretive connections. In Chapter 2, Chris Holcomb and Duncan A. Buell write about creating a first-year composition (FYC) corpus and what such a corpus reveals about the complexity of student writing. And in Chapter 3, Alexis Teagarden uses corpus analysis to understand the ways that student writers learn to use synthesis. Further, Teagarden suggests the importance of the “expert reader” (p. 55), who is instrumental in interpreting the results of machinic analyses.
Section 2, “Data Across Contexts,” contains four chapters that examine big data for purposes ranging from programmatic assessment, to genre analysis, to the measurement of transfer of FYC skills, to engagement in interdisciplinary writing. Laura Aull’s chapter is a standout: Aull traces discourse-level characteristics of student writing in directed self-placement essays. She argues that we should use such analyses to change our own practices of assignment design and “principled corpus building” (p. 97). And Kathryn Lambrecht's chapter shows how corpus analysis can be used to gain a better understanding of how writers engage in interdisciplinary work. In particular, Lambrecht uses big data to reflect on the difficulty of choosing and establishing a scholarly identity, particularly when that identity is affiliated with more than one set of disciplinary standards.
Section 3, “Data and the Discipline,” offers a deep discussion of disciplinary trends examined with big data methods. Chen Chen, for example, offers a corpus analysis of email subject lines from the disciplinary listserv WPA-L as a way to practice “disciplinographies” (p. 139). Chen is particularly interested in the listserv as a site of disciplinary knowledge production and argues that the discussions therein “help build and maintain a disciplinary community” (p. 149). Jenna Morton-Aiken combines archival theory with digital humanities in a chapter discussing relational architecture as an infrastructural methodology. Kate Pantelides and Derek Mueller use “data visualization, distant reading, and microanalysis” as a way to track disciplinary time on both macro and micro scales (p. 180). They use distant reading to construct a calendar of disciplinary conferences and then microanalysis to show how much time an individual researcher spends preparing for professional conferences. Finally, Cheryl E. Ball, Tarez Samra Graban, and Michelle Sidler discuss the potential of open data in writing studies, particularly the potential of creating “boutique data” and opening it to the public so that we might derive new insights.
Finally, Section 4, “Dealing With Data's Complications,” is the largest section and one of the strongest. Johanna Phelps responds to recent institutional review board (IRB) changes and explores creating disciplinary data sets for composition studies in order to both ease the burden of IRBs and create replicable, aggregable, data-driven research in the field. Next, Andrew Kulak thoroughly discusses cybersecurity and algorithmic accountability as methodological best practices. In particular, Kulak argues that we are responsible for attending to these concerns because of “our discipline’s connections with access, power, and knowledge as well as our historic relationship with marginalized, first-generation, and so-called nontraditional writing students” (p. 232) and offers several practical strategies for doing so. Juho Pääkkönen argues that interpretation plays a crucial role in model selection. Particularly, Pääkkönen wants researchers to understand that model selection is itself already an interpretation. Romeo García calls attention to the sets of stories we inherit in writing center research and how those stories can show us the “hauntings and hegemony” (p. 262) of the dominant culture. And Jill Dahlman writes about how to work with bad or missing data, using the example of inheriting old data as a writing program administrator. Dahlman offers solid advice about how to gain insight from data that might not seem relevant rather than throwing them out.
Overall, Composition and Big Data is a strong collection that asks difficult questions about how best to use big data methodologies and takes on the difficulties of humanizing data, storing and using data ethically, and questioning the interpretations we bring to the data through our methodological choices. But the book does have some weaknesses. Although each section contains many promising ideas, the section divisions themselves did not feel as useful to me. For example, Section 1 is titled “Data in Students' Hands,” but only the first chapter deals with students using data. So while the book does an admirable job centering ethics as a concern, some chapters fail to acknowledge Kulak’s concerns about disciplinary connections to “access, power, and knowledge” (p. 232).
In spite of such minor weaknesses, however, Licastro and Miller have compiled a useful disciplinary resource. In particular, this collection has a lot to say about big data and methodological design, offering many practical and theoretical questions for readers to consider in their own work. By having researchers write reflections on their own big data projects, Licastro and Miller have provided a model of scholarly metacognition. This approach shows us how to use metacognition to improve our work (whether that work is exploring archives, creating a corpus, or doing programmatic assessment). The book, then, will likely be especially relevant to those with interests in digital humanities and rhetorics and will provide a useful reference for graduate students seeking to more deeply understand the methodological choices at their disposal.
