Abstract
By design, randomized experiments (XPs) rule out bias from confounded selection of participants into conditions. Quasi-experiments (QEs) are often considered second-best because they do not share this benefit. However, when results from XPs are used to generalize causal impacts, the benefit from unconfounded selection into conditions may be offset by confounded selection into locations. This work shows that this tradeoff can lead to situations where estimates from QEs are less-biased from selection than are estimates from uncompromised XPs when drawing causal generalizations. This work establishes the conditions theoretically, demonstrates the idea empirically, and discusses the implications of the results.
It is natural for policy-makers and practitioners to ask about the potential for programs to achieve impact for an inference population that they are directly concerned with. Ideally, an uncompromised randomized experiment (XP) would be conducted in their specific context to answer the question decisively. However, if an XP is not possible, they may look to other sites, where the program has been used, to support an inference about the potential causal impact for their site.
Consider two options for doing this. The first, a more obvious and standard one, is to use an XP-based result from one or more study sites where the program has been evaluated—the “generalized from” site(s)—to infer impact for the “generalized to” site. This result may be adjusted for possible differences in the distribution of baseline characteristics that moderate the impact. This approach has the advantage that at the “generalized from” site, the result is not biased from confounded selection into conditions; however, because it involves a cross-site comparison of outcomes, the result may be biased from confounded selection into locations on characteristics that moderate the program effect (Hotz et al., 2005).
A second option is to make a cross-site comparison to infer just the missing outcome at the inference site. There are two possible situations for this depending on the missing outcome at the inference site. In the first scenario, the treatment has been used at the inference site, and the goal is to compare the performance given treatment at that site, to the performance in the absence of treatment using a comparison group from another site that has not used the program. This is the standard observationalist's application of a non-equivalent comparison group design (CGD), a type of quasi-experiment (QE) (Shadish et al., 2002). In the second scenario, the treatment has not been used at the inference site, and the goal is to compare the performance without treatment at that site, to the performance in the presence of treatment using a group from another site that has used the program. This is akin to a CGD, except it uses performance under assignment to treatment at another site to infer performance in the counterfactual condition for the group that has not received treatment at the inference site.
The second option has the advantage that it essentially uses an estimate of half the true impact result for the inference site (e.g., achievement in the presence of treatment for the first scenario, and performance in the absence of treatment for the second scenario); however, because this option involves a cross-site comparison of outcomes, it may be biased from confounded selection into locations on characteristics that either affect average achievement or that moderate program impact (Hotz et al., 2005).
The reader may well wonder why one would ever use the comparison group-based approach (the second option) in place of the XP-based one (the first option). The result based on an uncompromised XP has the advantage of being unbiased from selection at the study site, and therefore seems advantageous, even if it is achieved at a different site than the one for which the generalized impact is sought. Intuition suggests that adjusting results from an XP is the better option. In this work it is shown that, contrary to intuition, this is not always so. This finding is important because the less-biased option should be chosen whenever possible; therefore, it is critical to establish the conditions under which one approach yields less biased results than the other.
Clarifying the structure of this work will help the reader navigate it. There are five main sections. The first provides background. The second lays the groundwork for the method by describing three core scenarios investigated in the work. Each scenario offers a different solution to generalizing a causal effect of a program to an inference site using a comparison of outcomes from another site. The development of this section, especially the details of the first two scenarios, is important because the third scenario is a composite of the first two—it shows how a novel solution can be achieved by applying previous solutions in a new way. This section derives expressions for bias for the simplified scenarios of comparisons between just two sites. Both graphical and algebraic representations are given. 1 The basic example gets at the gist of the main idea—that generalized causal impact quantities, whether from experiments or non-experiments, can be biased for the inference site; and the levels of bias from each design can be quantified and compared. The third section sets up the empirical example by deriving expressions for average magnitudes of biases when generalizing impacts to multiple sites. The fourth section uses data from the Tennessee STAR (Student-Teacher Achievement Ratio) class size reduction multisite experimental evaluation to demonstrate an empirical approach to evaluating the alternative cross-site generalization approaches represented by the three scenarios. The fifth section concludes the work.
Background
Recently there has been a groundswell in advances in methods of evaluation for generalizing impact findings from experiments. The starting point of these methods is an impact finding from an XP conducted at one or more study sites. The generalization step involves adjusting the internally valid XP-based impact finding so that it reflects the local conditions for the inference population. Reweighting is a chief example (Schochet et al., 2014). The approach adjusts for differences between the study sample and the inference population in the distribution of moderators of impact. 2 Adjustments may also be made in terms of an index that summarizes the effects of multiple moderators on one dimension. Subclassification methods (e.g., Tipton, 2013) are an example.
Such XP-based approaches reflect an internal validity-first orientation: the starting point is an estimate of average impact from a true experiment, which is assumed to be internally valid. External validity—which addresses the extent to which a causal relationship holds over variations in persons, settings, outcomes or treatment variants (Shadish et al., 2002, p. 256)—is achieved after, and is based on the impact finding from a completed XP. This is consistent with a specific orientation in program impact evaluations that recognizes internal validity as the sine qua non (literally, “without which there is nothing”). The prioritization of internal validity makes establishing the causal nature of the relation under study the foremost concern—“the basic minimum without which any experiment is uninterpretable” (Campbell & Stanley, 1963, p. 5, in Shadish et al., 2002, p. 97). The implication is that when internal validity is of primary interest, then the priority is to first establish the causal relationship between variables, and only then address questions about the reach of the causal inference to contexts beyond the study.
The XP-based approach, however, is not the only one. Consider the case where a principal has implemented a program school-wide and would like to know if the program had an average positive impact on student achievement at her school. With the XP-based approach to generalization described above, she would look for an average impact finding from an XP conducted elsewhere, preferably from a locale similar to hers. She might then reweight the impact result from the remote site to more-closely reflect the distribution of characteristics of persons at her site. Alternatively, using a quasi-experimental (QE-based) approach, she may estimate impact for her site by comparing the average performance at her site, in which everyone has received treatment, to contemporaneous performance from one or more similar locales where the program has not been used. This is a standard non-equivalent comparison design (Shadish et al., 2002), with the inference locale being the treated site (i.e., the comparison is with a sample that provides a plausible counterfactual value for what the performance at the inference site would have been in the absence of treatment).
The comparison group-based approach accommodates also an alternative “reversed” scenario, where the program has not been used at the principal's (inference) site, and where she would like to know about the potential impact that could be achieved for her site. In this situation, the principal can again use the XP-based result, or, applying the comparison group-based strategy, she can infer impact to her site by comparing average performance at other similar locales, where the program has been used, against the average performance (in the absence of treatment) at her site. This approach is a version of the non-equivalent comparison design (Shadish et al., 2002), but where the comparison is made with a treated group to support a causal inference for the untreated site (i.e., the comparison is with a sample that provides a plausible counterfactual value for what the performance at the inference site would have been if the treatment had been used at that site).
For either of these scenarios, the XP-based result has the advantage that it is unbiased for the source sample (assuming the XP has not been compromised in some way). The success of generalization depends on completely identifying and adjusting for the effects of factors that produce a difference in impact between the experimental and inference sites (Cole & Stuart, 2010; Hotz et al., 2005; Imai et al., 2008). On the other hand, the comparison-group based approach has the advantage that it uses information from the actual inference site; however, the comparison with another site puts the result at risk of being biased from selection on confounders that affect average achievement, or on moderators that affect achievement by way of their interactions with the treatment. Put another way, the XP-based option uses an estimate of the true (unbiased) result, but it is from someplace other than the inference site; whereas the comparison-group based approach makes use of an estimate of the true outcome for the inference site itself; however, it is observed for just one condition (i.e., it is half of the unbiased solution for the inference location, and the other half must be inferred from someplace else).
Based on these scenarios, is not a priori clear that the XP-based generalization always yields a better (more accurate) result than the QE-based alternative. There is the prospect that under certain conditions, the QE strategy is the better option. This work addresses the following questions: (1) What are the expressions for bias for the XP- and QE-based generalizations described above? (2) Under which conditions is net bias in the XP-based generalization lower than in the QE-based result, and vice versa? (3) In a single empirical application, what are estimated levels of each type of bias, prior to and after applying site-level covariate adjustments?
The questions raised and explored in this work are motivated by an alternative perspective on the relationship between internal and external validity. Rather than seeing internal validity as the sine qua non, it treats the two forms of validity as interdependent. There is precedent for this in the program evaluation literature. Shadish et al. (2002), although accepting that internal validity is the sine qua non, also stressed that it is inseparable from external validity. They considered the latter as “the desideratum” (the purpose or objective) of educational research (Shadish et al., 2002, p. 97). The two kinds of validity— internal and external—are deeply complementary (Shadish et al., 2002). Further, they also interpreted the sine qua non of internal validity “in a strong sense,” that pertains to the situation where an experimenter has the option to exercise choice among several validity types. In that situation, internal validity may be prioritized at one stage of research, and external validity at another. Taking a stronger position, some evaluators outright rejected the precedence of one form of validity over the other, instead emphasizing the indispensability of one form of validity to another, and the role of contextual factors as co-causes of the results (e.g., Cronbach, 1975, 1982; Scriven, 2008). Cronbach, for example, stressed the “limited reach” of internal validity. For him identification of “the cause” of an effect has limited value without understanding the conditions and scope of that effect. Put another way, Cronbach considered that “if observations are not generalizable, then causal validity is irrelevant” (Albright & Malloy, 2000, p. 338). According to this interpretation, what we observe as the marginal impact in an XP-based evaluation, is the product of both the experimental manipulation plus the interactions of treatment with observed and unobserved moderators of the effect for that specific context. Any generalization requires making additional strong assumptions about the role and impact of those factors in the “generalized-to” context.
This work provides an alternative perspective on the positioning of internal and external validity. It shows that when causal generalization is the main concern, under certain plausible conditions, quasi-experimental comparisons can yield generalized inferences that are less-biased than XP-based ones. We develop the idea that both the XP- and QE-based inferences have potential for bias—either from confounders that affect average performance across sites, or from moderators that contribute impact heterogeneity across sites. Net bias in generalized impacts from QE-based designs depends on whether effects of confounders and moderators compound or offset one another. QE-based inferences may be preferred especially when this offset occurs because it can result in a cancellation of biases. An implication of this is that when causal generalization is the concern, the role of factors affecting internal validity (e.g., confounders) and external validity (e.g., moderators of impact) must be considered simultaneously. This point can be made theoretically, and studied empirically through so-called Within Study Comparisons (WSC) methods. We explore both in the sections that follow. 3
Three Scenarios to Motivate the Main Ideas of This Work
Scenario 1: Establishing the Experimental Benchmark
The goal is to estimate the impact of a program T relative to counterfactual C on outcome Y at a given site N. Express the true average impact quantity for this site as follows (XP stands for “experimental” to denote a quantity that would be estimated without bias through an uncompromised randomized experiment):

Representation of the experimental (XP) average treatment effect of assignment to T relative to C at site N.
The problem addressed in this work arises when information is missing about performance in one of the conditions at the inference site, N, thereby requiring the use of information from elsewhere (i.e., from other sites) to generalize impact to the inference site. These cases are explored in the next two scenarios.
Scenario 2: Inferring Average Impact at Site N When Treatment Is Provided to Everyone at the Site
In Scenario 2 the goal is to infer average impact for site N,
This situation would arise if a program is being implemented school-wide. The administrator may want to know if the program achieves positive average impact on student outcomes, compared to if it had not been used.
For this scenario, two options for inferring average impact at N are considered.

Representation of the experimental (XP) impact obtained at site M,
The difference between the true average impact at M, and the one at inference site N, is the bias in the former when used to generalize impact to the latter site. This is shown as
Lack of

Representation of the quasi-experimental average treatment effect (QE1) inferred to site N through comparison with site M.
The difference between the QE result,
Lack of
Option 1 uses the uncompromised XP-based impact quantity from site M,
Option 2 makes use of a quasi-experimental comparison of average performance between the treated sample at inference site N, and average performance in the absence of treatment at the comparison site
This is like Scenario 2, except no one at the inference site, N, has received the treatment. Scenario 3 may seem to be more-obviously about generalization, in the sense that an externally valid causal inference is being sought for a population that has not yet received the program or been involved in an experiment of the program. It is assumed that a randomized experiment cannot be conducted at N in the short term, and a plausible value for impact at N is needed immediately—perhaps to guide programming decisions for which an answer is urgently needed—and before a randomized experiment can happen.
As with Scenario 2, two options are considered to infer impact at N.
Figure 4 displays both the quantity of interest, which is the average impact for the inference site N,

Representation of the quasi-experimental average treatment effect (QE2) inferred to site N through a comparison with site M.
To formulate bias in
Option 1 uses the uncompromised XP-based impact quantity from site
Option 2 makes use of a quasi-experimental comparison of average performance between the treated sample at site M, and average performance in the absence of treatment at the inference site
This question is addressed in the next section. The alternatives discussed above are summarized in Table 1.
Scenarios for Inferring the Average Causal Impact of Program T to Site
Clarifying the Conditions Under Which a Quasi-Experimental Generalization Is Preferable to an Experiment-Based One
Each of Scenarios 2 and 3 explores two options for inferring impact at N, and the two scenarios suggest different rules for deciding the preferred alternative.
Scenario 2
When everyone at N (the inference site) receives treatment, impact may be inferred using
Scenario 3
When no one at N (the inference site) has received treatment, impact may be inferred using
Before moving to an empirical exploration of the questions, a further numerical comparison of the two options clarifies the conditions under which each should be preferred. Applying the criterion of lower net bias,
Figure 5 shows the relations graphically. It displays the space in which

Mapping the region over which
Before further discussing the alternatives, it is important to consider whether satisfying either of these conditions is even plausible. First, can the conditions in equation (11) be satisfied for certain plausible values of bias? Past empirical work shows that without adjustment for effects of covariates,
Second, given that the advantage of
Having established the condition under which
Method
Setting Up the Empirical Example
The empirical example of this work extends the methodology of WSC studies to empirically evaluate the levels of each type of bias discussed above. Traditionally, WSC studies are used to evaluate bias in non-experimental estimates relative to experimental benchmarks (pioneering studies are by Lalonde, 1986, and Fraker & Maynard, 1987).
The starting point for a WSC study is typically the impact finding from an uncompromised randomized experiment. This result serves as an unbiased benchmark quantity. A non-experimental result is generated by substituting the outcomes from the experimental control group with those from a different comparison group. The measured change in impact that results from this substitution estimates bias in the quasi-experimental comparison (corresponding to
Recently, WSC methods and other similar approaches have been extended to empirically evaluate discrepancies in generalized causal inferences from benchmark experimental impacts (Dehejia et al., 2021; Jaciw, 2016; Jaciw et al., 2022; Kern et al., 2016; Orr et al., 2019). The approaches parallel the standard one described above—the benchmark impact estimate from an uncompromised experiment at a site is replaced by a generalized impact quantity based on XP-based findings from one or more other sites. The difference between the generalized impact and the benchmark impact for the site reflects bias in the former quantity (corresponding to
The current work uses the “multisite variant” of the WSC method noted above (Bloom et al., 2005; Michalopoulos et al., 2004; Wilde & Hollister, 2007). With this approach, each site in a multisite experiment has a benchmark true value that may be estimated without bias through an uncompromised XP (i.e., each can play the role of inference site
Going forward, it is important to keep in mind that each WSC study is just one of potentially many. Multiple WSC studies are necessary to establish the empirical distributions of bias. That is, no single WSC study yields results that are definitive for understanding general conditions for bias. The accumulation of results from many studies provides the opportunity to develop research summaries of the findings, including specific procedural rules for avoiding or lowering bias that have been shown to work across many WSC evaluations. Summaries of this kind in traditional WSC evaluations include works by Bloom et al. (2005), Cook et al. (2008), and Glazerman et al. (2003). The results from the current study should therefore be considered as providing one piece of the evidence for evaluating how discrepant XP-based and QE-based generalizations are from the XP benchmarks for individual sites. Multiple similar efforts will help to consolidate what is known about conditions for bias more generally. The current work should also be considered a proof of concept, because it is a novel application of WSC methods. 11
Steps in the Method
Before describing the three steps of the method, one modification is presented. The causal generalizations discussed so far have been between just two sites, the inference (generalized to) site N, and the comparison (generalized from) site M. Going forward, instead of comparisons of outcomes being drawn with just one other site (
Site-specific biases in
Step 3. Summarize Bias for Each Alternative (
and
) Using Means of Squared Differences (i.e., Mean Squared Bias [MSB])
If overall bias was summarized by simply averaging over site-specific biases, then cancellation of positive and negative values would result in the underestimation of the average magnitude of bias (Bloom et al., 2005). Summarizing average levels of bias using the mean squared bias is one way to avoid this problem:
For
For
For
These expressions are convenient for addressing the main questions of this work, including about the degree of bias in comparison-group-based and XP-based generalizations when compared to experimental benchmarks. It is expected that on average the magnitude of bias in
Step 4. Adjust for Effects of Confounders and Moderators
As with standard WSC studies, a salient question is whether quantities that summarize bias, in this case the values of
Identifying Sources of Variation in Outcomes Across Sites
Estimates of
Using the multisite experiment version of a WSC design, what are the estimated average magnitudes of the discrepancies between site (benchmark) impacts and the three different versions of generalized impact corresponding to How do the magnitudes of the RMSB quantities in (1) compare to each other?
These questions will be addressed prior to and after covariate adjustments, and with and without adjusting for effects of class-level random sampling error. Given the focus of this work, the contrast of main interest is between estimates of Root Means Squared Bias for the experiment-based generalization,
Estimation
Hierarchical Linear Models (HLMs) (Raudenbush & Bryk, 2002) are used to obtain the following estimates:
Application
The application uses data from the Tennessee STAR (Student-Teacher Achievement Ratio) class size reduction multisite trial. Results from the study are reported by Finn and Achilles (1990), Mosteller (1995), and Nye et al. (1999, 2000, 2001, 2002). The multisite trial started in 1985 and lasted 4 years. In the original study, 6,400 kindergarten students and their teachers were randomly assigned to (1) small classes (13–17 students), (2) regular classes (22–25 students), or (3) regular classes with an aide within each of 79 schools (sites). Approximately 100 classes were allotted to each arm of the trial. The intervention continued through third grade. Teachers were randomly assigned to conditions within grades as the student cohort moved from kindergarten through third grade. The design aimed to retain students in the condition to which they were originally assigned over that period. Students who joined the study in intervening years were randomly assigned to conditions. In previous studies, regular classes, with or without an aide, are considered the control group, and this approach is adopted in this work. Having multiple control classes per school, allows estimation of sampling error for intermediate (class) units, which helps to address the question of the effect of ignoring the teacher level in estimates of MSB. 14
In the original study, outcomes were assessed on reading and math in kindergarten through third grade using the SAT-7 assessment. For the demonstration of this work, impacts are evaluated on second grade reading outcomes. The sample consists of students with posttests, who joined the study either in kindergarten or in first grade, and who remained in their assigned condition through second grade. To maintain a hierarchically structured dataset (i.e., with students nested in schools) analysis is limited to students who remained in the same school through second grade. Applying these criteria, the final dataset consists of 3,452 students among 314 second grade classrooms, among 73 schools. 15
The estimates of RMSB reported below are obtained: (a) prior to covariate adjustments; (b) after adjusting for effects of student-level school-centered variables (gender, eligibility for Free or Reduced Price Lunch, minority [non-White] status, years of experience teaching by the student's teacher, whether a student's teacher holds a Master's degree or higher, and end-of-kindergarten scores in math and reading, and the interactions of these covariates with treatment); and (c) after adjusting for the effects of both the variables in (b) and site-level (i.e., macro) variables, including site averages of uncentered student-level variables, and school urbanicity (i.e., whether a school is inner-city, suburban, rural or urban) and their interactions with treatment. 16
Results
The main results are displayed in Table 2 and Figures 6a and 6b. Each triplet of bars shows, from left-to right:
Estimates of Root Mean Squared Biases for Evaluating Several Approaches to Generalization.
Note. The values displayed are the square roots of estimates of corresponding MSB divided by the standard deviation of the outcome variable. Expression in the metric of the standardized effect size allows comparison with values for average impact and yearly expected growth. For example, the average impact of small classes on second grade reading performance in this experiment is .24-.25 (similar to results in Nye, Hedges and Konstantopoulos, 2000). For reference, annual expected growth in second grade reading scores is approximately .60 standard deviations (Hill et al., 2008).
Covariates at the student level (all school centered) are: gender (1=male, 0=female), eligibility for Free or Reduced Price Lunch (1=eligible, 0=non-eligible), and minority (non-White) status (1=minority, 0=non-minority), the years of teaching experience of a student's teacher, whether the student's teacher holds a Master's degree or higher (1=yes, 0=no), and end of kindergarten scores on tests of math and reading. Covariates at the school level are school averages of uncentered student-level covariates and variables indicating school urbanicity (whether a school is inner-city, suburban, rural or urban.)
Models include the main effects of the covariates and their interactions with treatment.
****p<.001, ***p<.01, **p<.05, *p<.10.

(a) Estimates of Root Mean Squared Bias without adjustment for class-level random effects (expressed in units of the standard deviation of the outcome distribution). (b) Estimates of Root Mean Squared Bias with adjustment for class-level random effects (expressed in units of the standard deviation of the outcome distribution).
What Is Observed? How Large on Average are the Discrepancies of Generalized Quantities From Experimental Benchmarks?
What Are the Main Takeaways From the Results?
The main takeaway points are separated into two types. The first address specifically results from Project STAR. The second address general implications for WSC studies of this kind. The focus is primarily on the difference between XP and QE2.
One conclusion from the derivations of this work is that when there is no opportunity for covariate adjustments and no information about the relationship between
However, under different situations and with more information, the preference could change. Consider first if there is a positive correlation between
Addressing Possible Limitations of the Empirical Application
In the STAR experiment, both students and teachers were randomly assigned to conditions, making the inclusion of the class random effect sensible; however, with multisite trials, modeling the intermediate level may be important even under different randomization schemes. For instance, if students are randomly assigned to conditions, but teachers are not, then bias may result from teachers’ selecting into conditions (e.g., certain teachers may jockey for the position to teacher students assigned to treatment). Even if randomization is blocked on classes, resulting in balance between conditions on student characteristics within classes, if treatment interacts with class averages of students attributes (or other class-level characteristics) this will be reflected in heterogeneity in impact among classes within sites. Unless these within-school effects are modeled explicitly, they may be misinterpreted as reflecting variability of impact across schools.
Results showed that the relative changes in estimates in school-level variance components were very similar when adding the same sets of school-level covariates to less- or more-parameterized models. That is, inclusion of specific sets of school-level covariates produced similar relative changes in estimated variance components regardless of how many school-level covariates were already in the model. If the models were overfitted, one would expect instead that changes in estimates of variance components from inclusion of additional school-level covariates to depend greatly on how saturated the model already is. The details of these tests are provided in Supplement A.
Continuing Efforts
Discussion
In their discussion of the prioritization of types of validity, Shadish et al. (2002) emphasize internal validity, as the sine qua non, in the context where “the experimenter can exercise choice within an experiment about how much priority to give each validity type” (p. 100) (with four types being considered: statistical conclusions, construct, internal, and external validity). They note further that “across a program of research, all validity types are a high priority” (p. 102). The approach discussed in the current work may be interpreted as putting internal and external validity on an even keel.
This work may also be seen as making a stronger claim: in application, causal generalization requires consideration of problems of internal and external validity as one, at least in the context of WSC studies. This work presented two options for drawing causal generalizations, XP, subject to
Put somewhat differently, a key idea of this work is that when causal generalization is the goal, bias from confounded selection on factors that affect average achievement in the absence of treatment (factors resulting in
A central idea of this work is that QE2, expressed as
The quantity QE2 may seem counter intuitive. An attendee at a conference where this work was presented objected to the possibility that bias in
Another possible concern is that researchers would never consider a site in which a program has not been implemented as the starting point for estimating impact. The normal course for an inference is to either run an experiment, or if a program has been used at the site, to seek a matched comparison group to establish a counterfactual outcome to persons who are receiving treatment (a standard QE). This work asserts that QE2 is valid, and prioritization of this causal quantity is in keeping with the idea that preferences for validity types are situated, and dependent on the aims of researchers and the maturity of the research program (Shadish et al., 2002). QE2 situates causal validity directly in the practitioner's world, where one must look to either experimental results from studies conducted somewhere else (e.g., through WWC), or a result that uses half the true solution for impact at the inference site (i.e., performance in the absence of treatment) and seeks a matched comparison group to establish a counterfactual outcome to persons who have not received treatment; that is, to infer how the inference group would have performed had they been assigned to treatment. This is the QE2 alternative.
Building on the idea of the prioritization of validity types that is described by Shadish et al. (2002), what are the implications of the approaches developed in this work for how validity types are prioritized? This work shows that the approach to establishing external validity that is dominant in experimental research, namely, run experimental study somewhere to obtain an internally valid causal result → adjust experimental result to support the external validity of a causal generalizations to a specific inference population, is not the only option for establishing a causal generalization. This order, which clearly prioritizes internal validity over external validity, may be the preferred approach for pure researchers who have time to amass a series of internally valid results before exploring the boundary conditions for effects. The alternative casual quantities compared in this work recognize that, in contrast, decision-makers on the ground may need causal evidence that applies directly to their specific context, and that prioritizes the use of performance outcomes from their site in order to increase relevance. The QE2 quantity may seem as an unfamiliar and hard-to-motivate alternative for some, but one much closer to evaluation purposes and relevant to decision points for local decision-makers.
Supplemental Material
sj-docx-1-aje-10.1177_10982140241246208 - Supplemental material for Hold the Bets! Should Quasi-Experiments Be Preferred to True Experiments When Causal Generalization Is the Goal?
Supplemental material, sj-docx-1-aje-10.1177_10982140241246208 for Hold the Bets! Should Quasi-Experiments Be Preferred to True Experiments When Causal Generalization Is the Goal? by Andrew P. Jaciw in American Journal of Evaluation
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
Appendix
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
