Abstract

The Experimental Methodology (EM) Section celebrates research that assists evaluators seeking to establish and describe causal relationships. In the face of several complex design-related questions, evaluators need both empirical and practical guidance to make informed choices on how to define treatment and comparison groups, account for validity threats in unique cases, navigate issues of complex implementation, and interpret causal findings. The EM Section strives to provide insight and clarity to support evaluators in making and defending these decisions. It also challenges scholars and practitioners to ask and answer related questions through thoughtful comparisons and empirical research.
Four articles appear in this issue's EM Section. Each considers separate, yet interrelated aspects of causal evaluation practice including describing, defining, and executing causal studies. Although each article addresses a unique set of questions, fundamentally, each study also has important implications for evaluators who face choosing between experimental and quasi-experimental designs.
The very nature of this choice may sound strange to some readers, because one's design is typically dictated by the specific research question(s) asked. However, this is not necessarily the case in evaluation (Wanzer, 2021), where design decisions will also depend on a host of contextual constraints. Although evaluators may have little input on whether to conduct an experiment or a quasi-experiment because of the evaluation sponsor's perceived ethical or feasibility concerns involving social experiments, with proper political navigation this need not be the case (Bell & Peck, 2016). Design decisions, therefore, must balance accuracy with other contextual needs. Accepting this reality, our motivating choice may instead be rephrased into a question: When is an experiment preferred to a quasi-experiment? To answer this requires weighing the tradeoffs among a host of constraining conditions. Important inputs to consider when answering this question relate to whether and under what conditions a quasi-experiment replicates an experiment. A larger body of literature involving within-study comparison (WSC) designs, also known as design replication studies, seeks to generate that evidence and blurs the boundaries that define the hierarchy of experiments and certain quasi-experiments (e.g., Wong et al., 2018).
One challenge to weighing tradeoffs between experiments and quasi-experiments lies in the terminology we use to describe designs. This is not specific to causal research but applies to all evaluation practices. Variations of designs emerge from methodological advances across disciplines where terminology may be discipline- or context-specific. In some cases, existing terminology is not easily adaptable to emergent research methods that come from other disciplines. Given the various contexts within which methodologists work, similar designs can generate different naming conventions or adopt practice-based names for designs that contradict use in other disciplines. Ironically, for causal research—a field that takes pride in precision—clear and consistent terminology is lacking.
The first article featured in the EM Section—titled Control, Exogeneity, and Directness: Understanding and Designing Quasi- and Natural Experiments by Dahlia Remler and Gregg Van Ryzin—takes on this very issue of the terminology used in describing causal evaluation designs. The authors’ efforts to examine assignment processes help to describe and clearly distinguish various forms of experiments and quasi-experiments that have evolved into inconsistent descriptions across various disciplines over the years. As part of their work, the authors include a helpful typology of eight causal study designs based on distinguishing factors of their design and structure. Also included is an excellent resource that shows how existing research examples, identified from the literature base, map onto each category to highlight the nuances of each design. Remler and Van Ryzin's work helps to provide much needed clarity on how we describe evaluation designs. I anticipate this work will be an excellent resource for teaching about causal designs as well as for serving to enhance review studies of causal evaluations. In sum, their typology will enhance the usability of causal evaluations and serve as an important reference for deciding between an experiment or a quasi-experiment.
Clarity in language can help us to better define the question of when an experiment is preferred to a quasi-experiment, but we also need rigorous methods to compare these designs. WSC researchers seek to articulate specific instances where certain quasi-experimental designs replicate the results of experiments. This work traces back to LaLonde (1986) and includes the influential efforts of Dehejia and Wahba (1999), which empirically extended upon LaLonde's ideas (Imbens & Xu, 2024). Later, Shadish et al. (2002) formalized that in the absence of random assignment, threats to validity can be carefully managed and even mitigated or relegated to nonnegligible through specific design choices. This, in turn, set the stage for promoting, under ideal conditions, certain “strong” quasi-experiments, for use in causal explanation. More recently Cook and colleagues have championed the examination of several different types of quasi-experimental designs using WSCs to understand the instances when they reproduce experimental results (see, e.g., Cook et al., 2008; Shadish & Cook, 2008; Wong & Steiner, 2018). Although in some cases a WSC can occur naturally or even retrospectively (Unlu et al., 2021), the design comparison is generally accomplished through a simultaneous study involving experimental (randomized) treatment and control groups and a separate quasi-experimental design involving treatment and (nonrandomized) comparison conditions. In these studies, the researchers’ main concern in constructing the WSC is to support internal validity. In other words, the WSC seeks to show whether the treatment effect produced via a quasi-experiment is comparable to the treatment effect produced via an experiment; and the WSC literature seeks to establish the conditions under which this comparability occurs.
In the next article in this EM Section—Hold the Bets! Should Quasi-Experiments Be Preferred to True Experiments When Causal Generalization Is the Goal?—Andrew Jaciw dives into WSCs to examine the performance of experiments and quasi-experiments when the goal of the study is causal generalization (i.e., supporting external validity). The study carefully capitalizes on the subtle differences in validity preferences, noting the importance of external validity when considering causal generalization. Jaciw's emphasis on providing helpful visual depictions of design and sources of bias help make this challenging and thought-provoking article much more approachable. The absence of WSCs that prioritize causal generalization represents an important opportunity for future research that not only supports a context-specific argument for comparing experiments with quasi-experiments, but also illustrates the possibility of parody in the choice of experiments versus quasi-experiments, particularly in situations when the evaluator places a higher value on supporting external validity over internal validity to promote the generalizability of causal claims.
Recognizing specific challenges and barriers of experimental and quasi-experimental designs can further illuminate important considerations when deciding how best to answer causal questions. The remaining two articles in the EM Section approach this topic from either side of the design pendulum. The third article—Challenges and Adjustments in a Multisite School-Based Randomized Field Trial by Debbie Hahs-Vaughn, Christine Depies DeStefano, Christopher Charles, and Mary Little—examines practice-based challenges of implementing a multisite cluster randomized trial (CRT). In this article, the authors reflect on a recently completed multisite evaluation. The authors use their experience with the evaluation to articulate five specific challenges evaluators face when carrying out multisite evaluations, and they recommend strategies for how to address each of these challenges. The authors’ details of their experience provide a rich and clear description of each challenge evaluators must consider when planning their study. Failure to anticipate and account for these challenges could easily derail a large-scale evaluation. Their experience serves as a reminder for evaluators to fully brainstorm what can go wrong before starting.
Although this article does not intentionally seek to compare experiments and quasi-experiments, astute readers will recognize that each challenge described in the multisite CRT also applies to the quasi-experimental counterpart. One might even envision the noted challenges as separate criteria for evaluating design choice considerations.
Likewise, the last article in this issue's EM Section characterizes two challenges tied to a specific type of quasi-experiment, a regression discontinuity (RD) design. In their article—titled Analysis of Regression Discontinuity Designs with a Binary Moderating Variable—Jason Schoeneberger and Christopher Rhoads use simulations to compare two analysis strategies for estimating “subgroup effects.” Moderating variables represent preexisting traits within a study sample that might lead to differential impacts for each subgroup. In an RD design, treatment exposure is defined by a threshold (or cut-point) on a continuous variable (called the running variable), and so a moderating variable may be distributed unevenly around the threshold. Schoeneberger and Rhoads demonstrate that choices about estimation strategies depend on the size of the bandwidth (i.e., how much data above and below the threshold is included in the analysis) and the prevalence of the moderating conditions within this bandwidth. Although this study focuses specifically on comparing estimation models of an RD design, it is also relevant to the WSC conversation, with RD designs being one of the quasi-experiments shown to mimic experiments under ideal conditions (Cook et al., 2008). It is a helpful reminder that a comparison of causal designs integrates context-based decisions across all practice elements, including analysis.
In many cases, the evaluator's choice to use an experiment or quasi-experiment may seem obvious or may even be predetermined, but several criteria factor into the decision. The four articles featured here draw attention to some of the difficulties that causal researchers face when selecting a design. Each article provides valuable insight to inform this decision.
Footnotes
Acknowledgments
The author is grateful to Laura Peck and Kyle Cox for review and comment on this From the Section Editor Note.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
