Abstract

Corpus data and methods have enjoyed broad application in different genres of discourse studies in recent decades (e.g. Sun et al., 2021). The advancement of corpus linguistics in this digital humanities era has greatly facilitated social scientists to answer the call for an intensification of systematic quantitative considerations in discourse studies. The book under review can inspire any researcher of interest to incorporate advanced inferential statistics into corpus-based discourse studies to examine the linguistic variation and change through the paradigm of comparative approaches. In other words, for readers of Discourse Studies, the significant value may lie more in its methodological guidance.
This volume comprises a neat introduction and a total of 11 chapters, which can be divided into four theme-based parts, namely, Corpus Dimensions and the Viability of Methodological Approaches (Chapter 1–2), Selection, Calibration and Preparation of Corpus Data (Chapter 3–5), Perspectives on Multifactorial Methods (Chapter 6–9), and Applications of Classification-Based Approaches (Chapter 10–11). In the introduction, the editors briefly present the book’s motivation, signpost the scope, practicality of its supplementary materials, and reflections on future directions, all of which enable the reader to better understand the main contents in subsequent chapters.
Part 1 mainly touches upon the inevitable trade-off between the reliability of small but detailed corpora plus annotations and the appeal of very large digital text databases for linguistics. The first chapter by Lukas Sönning and Julia Schlüter first illustrates how ‘richly annotated corpora can be leveraged to carefully navigate our research efforts in the era of big data’ (p. 42), based on the contrastive study of the /h/-onsets in two standard reference corpora, BNC and COCA, and the multi-billion-word database of Google Books Ngrams. However, the research of Sabine Arndt-Lappe and Sebastian Hoffmann in revisiting the Principle of Rhythmic Alternation (PRA) further emphasises that ‘much of the evidence in prior research is rather abstract and comes from large corpora accessed via the orthographic route only’ (p. 4). It argues that a close-up analysis from a small database can also converge with and add to the existed findings. These two cases reveal that the analyses of small and large databases best complement each other to ensure adequacy and reliability at the same time.
Further, Part 2 focuses on data preparation in corpus-based discourse studies. Specifically, the work by Fabian Vetter (Chapter 3) in comparing the press editorials’ sections of three different countries’ political discourse in ICE pointed out that ‘the observed differences can in fact be laid at the door of different sampling strategies applied by corpus compilers’ (p. 5). The chapter argues that a higher granularity of sampling schemes should be reached for the consideration of the representativeness of text samples. In the next Chapter 4, Sean Wallis and Seth Mehl clarified the appropriateness of baselines in quantitative comparison (e.g. normalised frequencies). Lukas Sönning and Manfred Krug (Chapter 5) also highlighted the hierarchical structure of corpora in a certain internal consistency (e.g. spoken or written modes, (sub-)registers, age groups, social identities, etc.).
Part 3 and Part 4 extend the discussion from traditional corpus approaches above to the combination of more advanced statistical methods. In Chapter 6–9, different multifactorial models are first exemplified, such as generalised linear mixed-effects models and random forests (in Tobias Bernaisch’s chapter), linear regression models and classification trees (in Matthew Fahy, Jesse Egbert, Benedikt Szmrecsanyi and Douglas Biber’s chapter), logistic regression approaches (in Natalia Levshina’s chapter), and multidimensional scaling and distance-based visualisations (in Ole Schützler’s chapter). In Chapter 10–11, the techniques of computational linguistics and machine learning, ‘two overlapping disciplines that have potential for supplementing and advancing linguistic research’ (p. 8), are illustrated in the final two cases (Gerold Schneider’s and Volker Gast’s chapter).
To sum up, this volume can function as a reference to offer practical guidance on how discourse data can be used in both traditional corpus and advanced inferential-statistics approaches rather than a textbook to understand the ontological position of corpus linguistics in discourse studies. Although this book is not intentionally designed for discourse analysts, it would also be of great interest and help to readers who are willing to adopt corpus methodology in their discourse analysis. Throughout the book, its most significant contribution is to display ‘a selection that is representative of the wide range of approaches current in corpus linguistics today’ (p. 10). It not only helps readers understand the research conventions and trends in corpus linguistics but also sheds light on different ways analysts can put advanced statistical methods into practice in corpus-based discourse studies.
In my opinion, some minor limitations can also be seen. First, its target readers, especially for Part 3 and 4, might be only established researchers and it is quite challenging for beginners, even graduate students, to fully understand their research methodology. More specifically, the involvement of different mathematical concepts and models (e.g. Euclidean distance, Bayesian regression, Possessum length, etc.) in their methods puts forward certain requirements for readers’ statistical background and digital literacy. Moreover, as each chapter is contributed by a different scholar, the inconsistent structural layout may not be reader-friendly enough to follow the whole book. However, its valuable methodological guidance for future corpus-based discourse studies still far exceeds the limitations above.
