Abstract
Researchers have been concerned with internal and external validity for decades, and the discussion continues. The present proposal is that there are less important and more important senses in which one can interpret internal and external validity, and these can be integrated with a taxonomy that includes theoretical, auxiliary, statistical, and inferential assumptions. The integration sheds light on recent exchanges in the literature on validity and suggests that the vaunted internal–external validity trade-off is false for more important senses of internal and external validity; internal and external validity increase or decrease together when there is an emphasis on underlying theories. Finally, the integration implies the desirability of some changes in typical research advice and practice.
Keywords
Much has been said about internal validity, external validity, and how they differ.
1
Vazire et al. (2022) provided descriptions that are representative of the extensive literature. They described internal validity as “the validity of causal inferences” (p. 165) and external validity as
the validity of inferences about how the observed effect will generalize beyond the specific conditions of the study. This includes generalizations to other people (or populations of whatever unit was sampled, if not people), other stimuli, other operationalizations or research designs, and other settings or times. (p. 166)
There has been much recent concern with external validity, with Yarkoni (2022a) claiming there is a generalizability crisis. Another external-validity concern stems from the burgeoning area of investigations into the generalizability of data collected through online crowdsourcing platforms such as Amazon Mechanical Turk, including whether previous support (e.g., Redmiles et al., 2019) still holds, and some data suggest it does not (Tang et al., 2022). The authors of these articles used external validity in ways consistent with Vazire et al.
A possible reason for the recent emphasis on external validity is a perception that internal validity has been studied more extensively. Lesko et al. (2020) specifically claimed this to be an important reason for reviewing threats to external validity and finding ways of addressing them. Like other external-validity researchers, Lesko et al. used a characterization of external validity in line with that put forth by Vazire et al. An additional reason for much recent focus on external validity could be concerns about a possible reproducibility crisis (Open Science Collaboration, 2015), which can also be considered an external-validity problem (Gelman, 2018).
However, the present goal is not to focus on either internal or external validity but to consider them simultaneously. The thesis is that there is a need to clarify (a) what we mean, or what we should mean, by internal and external validity, despite the large literature pertaining to both; and (b) the different levels of assumptions to be brought into play on which internal and external validity largely depend. In addition, these two issues can be integrated. We can commence from a place where there is wide agreement, which might even be considered a truism, that because internal validity tends to be maximized via carefully controlled laboratory experiments (e.g., Campbell & Stanley, 1963/2015; Shadish et al., 2002), which are abstracted from social contexts, such abstraction decreases the likelihood that the findings thereby obtained would generalize to those social contexts. And a strong concern with including socially relevant contexts renders difficult traditional laboratory control. Thus, we arrive at the received view of internal and external validity; there is a trade-off whereby maximizing one occurs at the expense of the other. To exemplify the received view, consider the following quotations by authorities across decades:
“Both types of criteria are obviously important, even though they are frequently at odds in that features increasing one may jeopardize the other” (Campbell & Stanley, 1963/2015, p. 5).
“Many of the choices among alternatives within each domain involve trade-offs or dilemmas” (Brinberg & McGrath, 1985, p. 42).
“It is a well-known methodological truism that in almost all cases there will be a trade-off between internal validity and external validity” (Cartwright, 2007, p. 220).
“When designing experiments and interpreting findings, the tension between experimentation goals and validity becomes apparent: Experiments provide the most direct way for determining causal effects and test theories because they maximize control and internal validity by simplifying, isolating, and making tractable even the most complex phenomena (Manzi, 2012; Pearl & Mackenzie, 2018), but these concessions are made at the cost of reducing external validity or the generalizability of the findings” (Lin et al., 2021, p. 855)
However, there is an ambiguity with respect to both internal and external validity. Let us consider internal validity first. There are two senses in which the effect can be due to the putative cause. Sense 1 is that the effect in an experiment is due to the manipulations or, more generally, that variance in the dependent variable is due to variance in the independent variables. This is consistent with both traditional (e.g., Brinberg & McGrath, 1985; Campbell & Stanley, 1963/2015; Shadish et al., 2002) and recent (e.g., Lesko et al., 2020; Vazire et al., 2022) accounts. However, Sense 2 is that the effect is due to the theorized cause (i.e., that the effect happened for the theoretically correct reason). Mook’s (1983) highly cited piece focused on theory, as did a recent article by van Rooij and Baggio (2021), although neither distinguished between Sense 1 and Sense 2 internal validity. To exemplify the difference between Sense 1 and Sense 2 internal validity, consider the classic cognitive-dissonance experiment by Festinger and Carlsmith (1959). They paid participants a large or small amount of money to say that a boring task is interesting and later obtained private attitudes. Festinger and Carlsmith predicted, and found, that people in the high-payment condition would have less favorable attitudes than people in the low-payment condition. The theoretical reason is that high payment provides a sufficient justification for lying about the interestingness of the task, whereas low payment does not. Hence, participants in the low-payment condition should have experienced cognitive dissonance as a result of insufficient justification and changed their attitudes to reduce the dissonance. However, although Bem (1967, 1972) agreed that the manipulation caused the effect (Sense 1 internal validity), he disagreed that it was for the theorized reason (Sense 2 internal invalidity). Instead, Bem argued that the manipulation influenced self-perceptions that subsequently influenced real attitudes. If we suppose Festinger and Carlsmith are right, then the effect is due to the manipulation (Sense 1 internal validity) and a difference between conditions in cognitive dissonance (Sense 2 internal validity). In contrast, if we suppose Bem is right, then although Sense 1 internal validity remains, Sense 2 internal validity disappears because the effect works for the wrong theoretical reason (self-perception as opposed to cognitive dissonance).
There is also an ambiguity with respect to external validity. Sense 1 is that the external-validity goal is to generalize the finding (as emphasized by Vazire et al., 2022), whereas Sense 2 is that the external-validity goal is to generalize the theory. Consider again the literature on cognitive dissonance and imagine that a researcher replicates Festinger and Carlsmith (1959) in a different context. In addition, suppose a set of additional studies supports the premise that the original Festinger and Carlsmith effect really is due to cognitive dissonance in the original context and not to self-perception; however, in the other context, a set of additional studies supports the premise that the replicated effect is due to self-perception and not to cognitive dissonance. In that case, there would be Sense 1 external validity because the effect replicates in the other context, but Sense 2 external validity would be suspect because the effect replicates in the other context for the seemingly wrong theoretical reason.
There is an asymmetry with respect to the relationship between Sense 1 and 2 internal-validity perspectives and the relationship between Sense 1 and 2 external-validity perspectives. To illustrate, consider again the original research on cognitive dissonance and let us suppose a compelling reason to believe that the manipulation did not cause the finding, which would be a Sense 1 validity problem. In that case, the finding fails to provide strong support for either cognitive-dissonance theory or self-perception theory. Speaking more generally, if there is a Sense 1 internal-validity failure, Sense 2 internal validity is out of the question. Thus, Sense 1 internal validity can be considered a prerequisite for Sense 2 internal validity. Sense 1 internal validity is necessary but not sufficient for Sense 2 internal validity.
Moving to external validity, this necessary yet insufficient characterization no longer works. There are many reasons why an experiment that works in one context might not generalize to another context. If a finding obtained in one context fails to generalize to another context, this Sense 1 external-validity failure need not have any implications whatsoever for Sense 2 external validity. For example, if Festinger and Carlsmith (1959) fails to replicate in a culture in which it is impolite to say, even on a private questionnaire, that the experiment is uninteresting—a Sense 1 external-validity failure—it would be unreasonable to interpret this as a serious problem for cognitive-dissonance theory. Rather, it would be necessary to perform the experiment in a different way that circumvents the politeness issue. If such changes were to result in success in the other culture, this would support Sense 2 external validity even with Sense 1 external invalidity remaining. Thus, it is not true that Sense 1 external validity is a prerequisite, or necessary condition, for Sense 2 external validity. Therefore, although Sense 1 internal validity is a necessary but not sufficient condition for Sense 2 internal validity, Sense 1 external validity is not even a necessary condition for Sense 2 external validity. This asymmetry is new to the literature on internal and external validity.
For the sake of clarity, the previous description of the asymmetry concerning the relationship between Sense 1 and Sense 2 internal-validity perspectives and Sense 1 and Sense 2 external-validity perspectives was stated in dichotomous language. However, this is an oversimplification because different types of validity are not necessarily there or not there but rather may be exhibited to different degrees. There can be many gradations. Thus, responsible research involves continually evaluating and reevaluating the worth of theories and the experiments that are believed to militate for or against them, as new thinking, new data, or both, continue to come into play.
A consideration of the recent debate involving Yarkoni (2022a), 38 comments, and a response (Yarkoni, 2022b) exemplifies the usefulness of distinguishing Sense 1 and Sense 2 validity. According to Yarkoni (2022a) there are many reasons why results may not generalize, researchers have not found sufficient ways to address these reasons, and so there is a generalizability crisis. However, Yarkoni’s focus was on Sense 1 external validity (the focus was the generalizability of results) as opposed to Sense 2 external validity. Although many commentators retained a Sense 1 external-validity focus, several of them criticized Yarkoni for a lack of consideration of the role of theory, although these researchers did not distinguish between Sense 1 and Sense 2 validity (e.g., Davidson et al., 2022; Harris et al., 2022; Hensel et al., 2022; Lakens et al., 2022; Maniadis, 2022; Turner & Smaldino, 2022). However, making the distinction suggests that the seemingly irreconcilable positions can, perhaps, be reconciled with worth accruing to both. Specifically, Yarkoni’s argument works better from a Sense 1 external-validity perspective than from a Sense 2 external-validity perspective, whereas the criticisms work better from a Sense 2 validity perspective than from a Sense 1 validity perspective. Thus, it is possible for both sides to have some merit but from different validity perspectives. And returning to the received view, it is possible to consider that this view has considerable worth from a Sense 1 validity perspective but less so from a Sense 2 validity perspective. This last is because, as we shall see repeatedly, factors that increase Sense 2 internal validity also increase Sense 2 external validity and vice versa. This is a crucial point that has not previously been made.
A Yarkoni (2022b) assertion segues to the next section. He characterized the criticisms focusing on his ignoring theory “unhelpful” (p. 74) although not necessarily wrong. This is possibly because calls for theory are unhelpful unless there is a way to elucidate the levels of assumptions necessary to relate theories and findings. I turn to this issue now.
The TASI Taxonomy
Trafimow’s (2019a) TASI taxonomy addresses the fact that there are different levels of research and different sorts of assumptions needed to traverse the gaps between levels. It is convenient to commence at the level of theory. Researchers propose theories that include theoretical assumptions. However, theoretical assumptions contain nonobservational terms, thereby rendering direct theory tests impossible. It is necessary to make auxiliary assumptions to bring a theory to the level of an empirical hypothesis that contains observational terms (Duhem, 1914/1954; Hempel, 1965; Lakatos, 1970, 1978; Meehl, 1990, 1997; Quine, 1951; Trafimow, 2009, 2012, 2017). The necessity of auxiliary assumptions implies that empirical findings can be attributed not only to the theory but also to auxiliary assumptions. Although empirical hypotheses sometimes can be tested without invoking additional layers of assumptions, sometimes they cannot. As an example of the former, Halley’s prediction about the return of the comet that now bears his name was based on a combination of Newton’s theory and auxiliary assumptions that were supported when the comet returned at the specified time.
However, most psychology predictions necessitate additional layers of assumptions to obtain desired specification. Consider an empirical hypothesis that scores in the experimental condition will exceed scores in the control condition. This empirical hypothesis could be cashed out through means, medians, 75th percentile ranks, and other summary statistics. It is necessary to make statistical assumptions to bring an empirical hypothesis to further specification rendering a statistical hypothesis. Finally, if one wishes to perform an inferential statistical analysis, such as using the null-hypothesis significance testing procedure, Bayesian procedures, or others, it is necessary to add inferential assumptions to the mix to render an inferential hypothesis. One difference between a statistical hypothesis and an inferential hypothesis is that the former concerns sample statistics whereas the latter concerns population parameters. Hence, statistical and inferential assumptions can be the same; for example, one may need to assume normal distributions to justify framing the statistical hypothesis in terms of sample means, but also to justify framing the inferential hypothesis in terms of population means. But statistical and inferential assumptions can differ too. For example, the ubiquitous assumption of random selection from a population (Berk & Freedman, 2003; Hirschauer et al., 2020) is practically always necessary for inferential generalizations to populations but often not required for statistical hypotheses. Even with a lack of random selection, the sample statistics, whether they be means, correlation coefficients, and so on, can support or not support the empirical hypothesis. But lack of random selection would strongly decrease confidence that the sample statistics are good estimates of corresponding population parameters. That decreased confidence might, in turn, decrease researchers’ willingness to use the sample effect to draw conclusions about the empirical hypothesis or theory.
Most statistics textbooks emphasize that descriptive statistics concern samples and inferential statistics concern generalization to populations. Hays (1994) stated:
Although descriptive statistics forms an important basis for dealing with data, a major part of the theory of statistics is concerned with another question: How does one go beyond a given set of data and make general statements about the large body of potential observations, of which the data collected represent but a sample? (p. 1)
Amrhein et al. (2019) suggested that inferential statistics can be treated as descriptive statistics, as p values or other data-based statistics vary across samples, which suggests the possibility of collapsing statistical and inferential assumptions into a single category. However, most researchers and statisticians continue to consider descriptions of samples and inferences to populations fundamentally different processes but with some overlap because sample statistics are used to help draw conclusions about population parameters. In keeping with the traditional position, the TASI taxonomy treats statistical hypotheses about sample statistics as differing from inferential hypotheses about population parameters.
In summary, there are theoretical assumptions in a theory, auxiliary assumptions needed to bridge the gap from the theory to the empirical hypothesis, statistical assumptions needed to bridge the gap from the empirical hypothesis to the statistical hypothesis, and inferential assumptions needed to bridge the gap from the statistical hypothesis to the inferential hypothesis. In turn, a model can be denoted that includes only theoretical and auxiliary assumptions, exemplified by Halley’s comet, as a modelTA, where the subscripts refer to the classes of assumptions included in the model. Continuing, a model can be denoted that includes only theoretical, auxiliary, and statistical assumptions as a modelTAS. Finally, a model can be denoted that includes theoretical, auxiliary, statistical, and inferential assumptions as a modelTASI. Thus, there are four levels of theorizing or hypothesizing: theory, empirical hypothesis, statistical hypothesis, and inferential hypothesis. And there are four levels of assumptions needed for each: To reiterate, theoretical assumptions allow the theory, auxiliary assumptions traverse the distance from the theory to the empirical hypothesis, statistical assumptions traverse the distance from the empirical hypothesis to the statistical hypothesis, and inferential assumptions traverse the distance from the statistical hypothesis to the inferential hypothesis. Finally, there are three types of models: modelTA, modelTAS, and modelTASI. The TASI taxonomy is depicted in Figure 1.

Depiction of how adding auxiliary assumptions to a theory results in an empirical hypothesis, with the combination of theoretical and auxiliary assumptions denoted by modelTA; how adding statistical assumptions to the empirical hypothesis results in a statistical hypothesis based on theoretical, auxiliary, and statistical assumptions denoted by modelTAS; and how adding inferential assumptions to the statistical hypothesis results in an inferential hypothesis based on theoretical, auxiliary, statistical, and inferential assumptions denoted by modelTASI.
Implications of TASI for Internal and External Validity
One reason theories are valuable is that predictions can be derived from them, of course, in conjunction with assumptions at other levels in the TASI taxonomy. Theoretical value is, in part, supported by Sense 2 internal validity. But theoretical value is also supported by the extent to which theories apply outside the original research paradigm. In his famous defense of external invalidity, Mook (1983) stated clearly that having strong experimental evidence pertaining to the theory (Sense 2 internal validity) was much more important than having results generalize (Sense 1 external validity). Mook had little to say about Sense 2 external validity, but the present distinction addresses that lacuna. It seems incontrovertible that a theory is problematic if it works only in the context in which it is developed but nowhere else (Diener et al., 2022; Laudan, 1984). Lin et al. (2021) impugned strongly against the worth of theories that work only in the paradigms in which they were developed. A proviso that bears emphasis is that a fair test of Sense 2 external validity requires that auxiliary, statistical, and inferential assumptions are appropriate for the contexts to which one wishes to generalize.
Sense 2 internal and external validity strongly influence theoretical value, and a prerequisite for Sense 2 internal and external validity is having high-quality assumptions at different levels in the TASI taxonomy that can eventually be shown to withstand tough empirical tests. Without the TASI taxonomy, or something like it, there is no way to traverse the distance from nonobservational terms in theories to the specificity necessary for impressive empirical tests and hence no way to achieve Sense 2 internal and external validity. This, I contend, partly addresses Yarkoni’s (2022b) admonition that calls for more theory are unhelpful. The unhelpfulness is because of not addressing the categories of assumptions necessary to traverse the distance from theory to empirical, statistical, or inferential hypotheses.
Because theoretical value depends largely on Sense 2 internal and external validity, which, in turn, require negotiating levels of the TASI taxonomy, these are inextricably linked. However, the present exposition thus far lacks an explication of how Sense 1 and Sense 2 internal and external validity play out at different levels of the TASI taxonomy, to be addressed presently, with examples.
Theoretical and auxiliary assumptions (modelTA )
Consider the theory of reasoned action and its descendant theories (Ajzen & Fishbein, 1980; Fishbein, 1980; Fishbein & Ajzen, 1975, 2010), in which attitudes are assumed to cause behavioral intentions. This is a theoretical assumption, but it is not the only one. For example, the statement implies that attitudes exist, behavioral intentions exist, and there is a causal connection between them. And there are additional assumptions about other variables and their connections to attitudes and behavioral intentions, although a discussion of these would be superfluous to the present purposes. Any of the theoretical assumptions might be false, although all might be true.
When attempting to test the theory, or a particular aspect of the theory, such as the assumption that attitudes cause behavioral intentions, there is an immediate problem. “Attitudes” and “behavioral intentions” are nonobservational terms. Because attitudes and behavioral intentions cannot be observed, there is no way to provide a test unless something is added; there must be a way to add auxiliary assumptions to reduce the theoretical statement to an empirical hypothesis that contains observational terms. For instance, in an experimental study, a researcher might link unobservable attitudes to an observable essay designed to manipulate them under the expectation that the essay influences attitudes. The theory does not state that the attitude manipulation influences attitudes, so the assumptions necessary to link the manipulation to the attitude construct are auxiliary to the theory—they are auxiliary assumptions. Similarly, the theory does not state that the behavioral-intentions scale successfully measures behavioral intentions, so the assumptions the researcher makes linking the observable scale to behavioral intentions constitute yet more auxiliary assumptions. A common decree is that theoretical terms must be operationalized, but the decree may be unfortunate because it obscures the necessity to make auxiliary assumptions to traverse the distances between nonobservational theoretical terms and observational empirical terms. If the auxiliary assumptions that traverse the distances between nonobservational terms in theories and observational empirical terms in hypotheses are not stated or only stated vaguely, it is difficult to assess the extent to which the manipulations or measures are sound. Because explicit auxiliary-assumptions statements are rare in psychology, this is an area in which there is much room for improvement.
Auxiliary assumptions aid in traversing the distance between nonobservational terms in theories and observational terms in empirical hypotheses, but that is not their only function. Auxiliary assumptions can set initial conditions, such as that randomly assigning participants to essay versus no-essay conditions renders the two groups initially equivalent on all relevant factors (an assumption that is not necessarily true). Nor need auxiliary assumptions be explicit; it is usually an implicit auxiliary assumption that the research assistant distributes the correct forms to participants.
The combination of theoretical and auxiliary assumptions—that is, the modelTA—implies that the attitude manipulation should cause scores on the behavioral-intention scale to be larger in the essay condition than in the no-essay condition. There is an ambiguity in what larger scores mean that is addressed in detail in the subsequent section, but let us tolerate the ambiguity for now. Suppose that the researcher suffers an empirical defeat because scores on the scale are not larger in the experimental than control condition. The empirical defeat could be blamed on the theory being wrong or on one or more auxiliary assumptions being wrong (e.g., Lakatos, 1970, 1978). Alternatively, suppose the researcher enjoys an empirical victory; scores on the scale are larger in the experimental than control condition. The empirical victory could be credited to the theory or to the auxiliary assumptions (Trafimow, 2017). Researchers’ perceptions of the quality of the auxiliary assumptions, and perhaps additional empirical evidence of their quality, likely would influence their judgments about the extent to which to attribute empirical defeats or victories to the wrongness or rightness of theories, respectively.
Let us optimistically suppose that the auxiliary assumptions in our reasoned-action experiment are of high quality and the experiment always works. The attitude manipulation manipulates attitudes and does not manipulate other causally relevant variables, the behavioral-intention scale validly indexes behavioral intentions, the randomization is successful, the research assistant distributes the correct forms, and so on. Furthermore, the anticipated difference between conditions on the intention scale occurs. Under our optimism with respect to the auxiliary assumptions, internal validity in both the Sense 1 and Sense 2 interpretations increases. For example, our belief that the randomization process is successful increases our confidence that the manipulation causes the effect (Sense 1 internal validity). Our belief in the validity of the manipulation and measure increases our confidence that the manipulation causes the effect for the right theoretical reason (Sense 2 internal validity). Moreover, additional measures of attitudes and of competing constructs could increase our confidence in Sense 2 internal validity if the former showed the expected effect, with trivial effects on measures of competing constructs (Fiedler et al., 2021).
Moving to external validity, our belief that the attitude manipulation works, as it is supposed to work, might not increase confidence that the experiment would work if performed in another culture because that which influences attitudes in one culture might not influence attitudes in another culture (potential Sense 1 external invalidity). But it might increase confidence that the theory could be made to apply in another culture (Sense 2 external validity). In fact, the Sense 1 external invalidity, if it were to happen, could benefit Sense 2 internal and external validity simultaneously. To illustrate this surprising statement, suppose that manipulating attitudes in Culture A requires an essay that focuses on the individual, whereas manipulating attitudes in Culture B requires an essay that focuses on the collective. 2 In that case, there would be no reason to expect the experiment that works in Culture A to also work in Culture B because the manipulation misses the mark for experiments performed in Culture B. However, suppose that the experiment does work in Culture B when the auxiliary assumptions are changed in accordance with knowledge about that culture, so that the essay focuses on the in-group when the experiment is performed in Culture B. That the new experiment works in Culture B (Sense 2 external validity) whereas the replication of the original experiment does not work (Sense 1 external invalidity remains) supports the premise that the theory generalizes to Culture B under appropriate auxiliary assumptions. Not only is Sense 2 external validity supported, but Sense 2 internal validity is supported as well by dint of the demonstration that combining the theory with another set of auxiliary assumptions results in another empirical victory for the theory.
In terms of the received view of an internal–external validity trade-off, there are implications. It may well be true that maximizing Sense 1 internal validity decreases Sense 1 external validity if the original experiment were designed to work in the United States but not in another culture. Even here, the received view might not apply fully because there is no reason to believe, say, that successful randomization that aids Sense 1 internal validity necessarily harms Sense 1 external validity. With respect to Sense 2 internal and external validity, there is more clarity, although now in contradiction to the received view. That is, better auxiliary assumptions lead to better confidence in the theory (Sense 2 internal validity), which, in turn, leads to more confidence that the theory would generalize when implemented with appropriate auxiliary assumptions that work in the new culture (Sense 2 external validity). Moreover, consider again that the same theory might require different sets of auxiliary assumptions to work in different cultures, thereby forcing an internal–external validity trade-off at the Sense 1 level, whereas internal and external validity are both increased at the Sense 2 level. It is a fascinating paradox that internal and external validity can be in trade-off at the Sense 1 level but function in unison at the Sense 2 level. This paradox is soon explained on contemplation that a quality modelTA in each culture simultaneously benefits both internal validity and external validity at the Sense 2 level, whereas this may not be so at the Sense 1 level.
Mediation
A typical way to address the problem of auxiliary assumptions is through mediation. Suppose an independent variable is alleged to influence the dependent variable, but an auxiliary assumption of an intervening mediating variable (e.g., the focal construct) is necessary to logically connect the independent variable to the dependent variable. The researcher might conduct a mediation analysis, with a path diagram showing an arrow from the independent variable to the mediating variable and another arrow from the mediating variable to the dependent variable. There also might be a so-called direct path illustrated by an arrow directly from the independent variable to the dependent variable. Does establishing a path diagram, with reasonably sized path coefficients on the arrows, convincingly demonstrate that the researcher correctly combined the theory and auxiliary assumptions to render the empirical prediction? Not necessarily. In a purely correlational study, there are very many potential ways to interpret the data, with the touted way being only one of them. In fact, Kline (2015) showed that 18 potential models might be consistent with the data in the three-variable case. If the independent variable is manipulated, that constrains the possibilities but nevertheless leaves five still plausible.
Returning to the essay versus no-essay manipulation, assume a measure of attitude and of behavioral intention, both of which receive larger summary statistics in the essay than no-essay condition. A researcher who supported the theory of reasoned action would like to make the case that the manipulation influenced attitudes that subsequently influenced behavioral intentions. To support that desire, the researcher might perform a mediation analysis to show that there is a nontrivial path coefficient from the manipulation to the attitude measure and from the attitude measure to the behavioral-intention measure. The researcher might even try to show that the so-called direct path from the manipulation to behavioral intentions reduces substantially when attitudes are considered. However, suppose we have a wrong auxiliary assumption so that the attitude measure does not really measure attitude but rather affect. Nor is this a trivial possibility (see Fishbein, 1980). In that case, although the mediation analysis enhances Sense 1 internal validity, subject of course to the other path models that could be constructed to account equally well for the data (Kline, 2015), a Sense 2 internal-validity problem would remain.
Thus, we have seen that mediation evidence is weaker, even in the presence of a true experimental manipulation, than researchers typically appreciate. Many articles reinforce that mediation effects are susceptible to many explanations and consequently fail to provide definitive evidence for any single explanation (Fiedler et al., 2011; Grice et al., 2015; Kline, 2015; MacKinnon et al., 2000; Tate, 2015; Thoemmes, 2015; Trafimow, 2015). The point is not that researchers should not test mediation, only that such tests fall far short of being definitive.
Manipulation checks
Another way to test auxiliary assumptions is to perform manipulation checks (Fiedler et al., 2021). If a manipulation is supposed to influence a construct, with that influence subsequently transmitted to the dependent variable, a manipulation check can be helpful in supporting the premise that the manipulation really does influence the focal construct. A limitation is even given that the manipulation influences the focal construct, that need not indicate that the effect on the dependent variable is because of the effect on the focal construct. For instance, practically every attitude theory agrees, if conditions are right, that manipulating attitudes should influence a relevant behavior. Suppose a researcher performs an attitude manipulation, performs a manipulation check to show that the manipulation really does influence attitude, and shows, too, an effect on behavior. Although the manipulation check supports the premise that the manipulation influences attitude, it nevertheless remains possible that the manipulation also influences some other construct, such as affect or mood, and it is the latter construct that is responsible for the effect on the dependent variable.
There are ways to make progress. One way, discussed in the foregoing subsection, is to perform a mediation analysis. A potential problem, as we have seen, is that affect or mood is likely correlated with (a) the manipulation, (b), attitude, and (c) behavior, such that there is no way for the mediation analysis to rule out affect or mood as alternatively explanatory. A better way is to perform additional measurements such as measuring affect or mood. If there is attitude change but not affect or mood change with a subsequent effect on the dependent variable, then attitude is a better explanatory construct than affect or mood for the effect. In this scenario, the additional measurement may do little for Sense 1 internal validity, which is perhaps why additional measures of this type are underused, but the additional measurement greatly enhances Sense 2 internal validity. A caveat is that even this prescription is not foolproof. It is possible that the affect or mood measures are invalid, and so the lack of an effect need not conclusively militate in favor of attitudes over affect or mood as the reason for a change in the dependent variable. A way to address the caveat might be to perform additional studies to provide independent evidence for the validity and sensitivity of the affect or mood measures.
An example in this direction is provided by typical research on terror-management theory (J. Greenberg et al., 1992, 2000, 2001; T. Greenberg et al., 1994). Terror-management researchers wish to attribute their effects to differences in death thought accessibility. However, there may be differences in affect that constitute a potential alternative explanation. To reduce the plausibility of the alternative explanation, terror-management theory researchers routinely have participants complete the Positive and Negative Affect Schedule (or PANAS) to show that there are only trivial effects of their manipulations on affect, thereby reducing the plausibility of affect as a competing explanatory construct and enhancing Sense 2 internal validity. Of course, that terror-management theory researchers are usually able to convincingly eliminate affect as a plausible alternative explanation does not necessarily indicate a lack of other plausible alternative explanations or problems (McConnell, 2018).
Theoretical, auxiliary, and statistical assumptions (modelTAS )
Consider a theory of prejudice by Stephan and Stephan (2000), that perceived threat associated with an out-group causes prejudice toward that outgroup. To test the theory, suppose a researcher performs a threat-inducing manipulation in which participants in the experimental condition are made to feel more threat than participants in the control condition. The dependent variable is scores on a prejudice scale, where the hypothesis is that scores in the experimental condition will exceed scores in the control condition. But surely the researcher cannot mean that every score in the experimental condition will exceed every score in the control condition. More likely, the idea would be that the mean should be greater in the experimental than in the control condition. However, there might be some extreme scores, in which case the researcher might be better off proposing that the median score will be greater in the experimental than control condition.
Or perhaps the distribution is skew normal rather than the usual assumption of a normal distribution. Normal distributions have two parameters: mean μ and standard deviation σ. In contrast, skew-normal distributions have three parameters: location ξ, scale ω, and shape α. Azzalini (2014) and Azzalini and Capitanio (1999) provided skew-normal details, but for the present purposes it is merely necessary to quickly consider skew-normal distributions when the shape parameter equals zero or does not equal zero. When the shape parameter equals zero, the distribution is normal, the location equals the mean, and the scale equals the standard deviation. In symbols, when α = 0, ξ = μ, and ω = σ. However, when the shape parameter does not equal zero, then the location does not equal the mean, nor does the scale equal the standard deviation. In symbols, when α ≠0, ξ ≠ μ and ω ≠ σ. Thus, the family of normal distributions constitutes a special case of the family of skew-normal distributions. Because most distributions are skewed (Blanca et al., 2013; Ho & Yu, 2015; Micceri, 1989), the family of skew-normal distributions is more generally applicable than the family of normal distributions. Moreover, although differences in locations are often in the same direction as differences in means, it is possible for the two types of differences to be in opposite directions, thereby implying opposing substantive stories (Trafimow et al., 2019).
Continuing with the experiment, suppose that the means in the two conditions support the theory, whereas the locations in the two conditions are in the opposite direction, thereby contradicting the theory. Which difference should the researcher emphasize? Much depends on the statistical hypothesis and the underlying statistical assumptions. In the present case, in which the goal is to test the theory, it really is necessary to be able to argue that the manipulation shifts the experimental distribution relative to the control distribution, and so the difference in locations should be taken more seriously than the difference in means. However, if the goal were an applied goal, whether the distribution shifts might be of lesser importance and the probability of a higher prejudice score in the experimental than control condition might be of greater importance. In this latter case, the difference in means might be more indicative of the effect of the manipulation on the probability of higher prejudice scores in the experimental condition relative to the control condition. The lesson is not that locations are always superior to means, or the reverse, but rather that whether it is best to narrow the empirical hypothesis to a statistical hypothesis featuring a difference in means, locations, or something else depends not only on the statistical assumptions one makes but also on the researcher’s goals.
Well, then, let us optimistically suppose that our prejudice experiment generalizes across multiple sets of auxiliary assumptions, such as different threat-inducing manipulations and different prejudice scales (or if not, that the theory works under appropriate auxiliary assumptions for the study contexts). And suppose generalizability across multiple sets of statistical assumptions, whereby the findings support the prediction regardless of whether the statistical hypothesis is framed in terms of differences between means, locations, medians, and so on. In that case, generalizing across statistical assumptions provides confidence that the ability to support the prediction is not importantly influenced by the statistical assumptions chosen. Although generalizability is usually conceived of as supporting external validity, here we see that generalizability across varying statistical assumptions supports the premise that the findings also occur for the theoretically correct reason, and supporting evidence is not just a matter of choosing lucky statistical assumptions. In short, going this route supports the premise that (a) the experiment works for the correct theoretical reason and (b) generalization across different statistical assumptions. Of course, support falls well short of proof. No matter how much generalization there is across different versions of the modelTAS, a yet different version might (a) better reflect reality and (b) result in contrary evidence (Trafimow, 2019b).
Alternatively, we could be pessimistic and assume a problem at the level of statistical assumptions. For example, we might use means when we ought to be using locations, with differences in locations in the opposite direction of differences in means. In that case, we have an obvious external validity problem at the level of statistical assumptions. Although the difference in means supports that there is an effect and that it likely is due to the manipulation (Sense 1 internal validity remains), the more important difference in locations contradicts the theory. That is, we have a Sense 2 external-validity problem because the effect fails to generalize to the appropriate type of summary statistic. This Sense 2 external-validity problem is also a Sense 2 internal-validity problem because we have less confidence that even the difference in means is for the theoretically correct reason. Sense 2 internal and Sense 2 external validity go together because mishaps at the level of statistical assumptions decrease both.
Theoretical, auxiliary, statistical, and inferential assumptions (modelTASI )
Psychology researchers rarely stop with a statistical hypothesis but typically also wish to perform a test of an inferential hypothesis. For example, a researcher might presume the inferential hypothesis that the population means in the experimental and control conditions are the same, with the hope of obtaining a small p value to militate against it and in favor of an alternative inferential hypothesis that the population means are different. However, to traverse the distance between the statistical hypothesis pertaining to differences in sample summary statistics to the inferential hypothesis pertaining to differences in population summary statistics requires the addition of inferential assumptions. For example, some have complained that researchers practically never sample randomly from a defined population (Berk & Freedman, 2003; Hirschauer et al., 2020), although this issue may be mitigated in true experiments with random assignment to conditions depending on whether the goal is to generalize across randomizations (for the same participants) or outside the participants in the study. To accomplish the latter, it is necessary to sample randomly from the population unless there is a theoretical reason to support generalization to a population of potential participants. Aside from this issue, there are so many assumptions that Bradley and Brand (2016) and Trafimow (2019a) proposed assumption taxonomies.
That most researchers compute p values need not indicate that this is what they ought to do (Gelman, 2018). Some have argued in favor of Bayesian approaches (e.g., Gelman et al., 2013), second-generation p values (e.g., Blume et al., 2019), the a priori procedure (e.g., Trafimow et al., 2019), and others. For present purposes, it is unnecessary to engage the inferential debate. 3 But it is necessary to stress that no matter what inferential approach one uses, provided one distinguishes between sample statistics and inferences to corresponding population parameters, inferential assumptions are unavoidable (Bradley & Brand, 2016; Trafimow, 2019a). In turn, the quality of the inferential assumptions influences both internal and external validity.
To illustrate, consider again the prejudice experiment. Suppose that the sample mean in the experimental condition exceeds the sample mean in the control condition. The researcher almost certainly would not be willing to settle for the conclusion that the sample means differ but would wish to make a larger statement that the samples come from populations with means that differ. If the researcher is unable to support a claim that the difference in sample means signifies a corresponding difference in population means, that lack is multiply problematic. One problem is that there is no reason to believe the effect would replicate with alternative samples (a Sense 1 external-invalidity problem). Another problem is that there is no reason to believe that the theory generalizes to the populations of interest (a Sense 2 external-validity problem). In turn, these external-validity problems also compromise Sense 1 and Sense 2 internal validity. Lack of confidence in a population difference decreases confidence that the effect was due to the manipulation as opposed to getting lucky (Sense 1 internal invalidity) or that the effect is for the theoretically correct reason (Sense 2 internal invalidity). Thus, we have a reversal of an often-stated truism that internal validity is a prerequisite for external validity. We now see that external validity, under some circumstances, can be considered a prerequisite for internal validity.
Before proceeding to applied research, there is a potential source of confusion to be addressed pertaining to the necessity to think in terms of populations. 4 There is no such necessity provided the researcher has no desire to generalize to a population, in which case we revert to a modelTAS or perhaps even a modelTA. However, there is a price that might need to be paid, in typical psychology research, for eschewing populations. To exemplify this price, consider again the prejudice experiment, but under the stricture that the mean difference in prejudice between the experimental and control conditions has nothing to do with the population difference between the two conditions, whether this is a population of randomizations, a population of studies that could have been performed, or a population of people. In this hypothetical case, much is unclear, including that the sample reflects anything outside of what happened in the specific experiment that was conducted. Getting lucky would remain a plausible alternative explanation, thereby compromising Sense 1 internal validity. And because Sense 1 internal validity is a prerequisite for Sense 2 internal validity, the latter would also be compromised. A lack of justification for concluding that the prejudice difference would happen similarly if we were to use the population of concern renders difficult justifying that threat causes prejudice. One contribution of integrating the distinction between validity senses and the TASI taxonomy is the consequential realization that the ability to generalize to populations may be a prerequisite for Sense 1 and Sense 2 internal validity, a reversal of the cliché that internal validity is a prerequisite for external validity.
Applied Research
There is much recent applied research. 5 Applied research may be theory-based, but it does not have to be. For example, an intervention based on reasoned-action theory is theory-based and depends importantly on the ability to make quality auxiliary assumptions to bring nonobservational terms in the theory to the level of observational terms pertaining to the intervention. In this case, the full modelTASI is in play. However, applied research need not be theory-based, which poses a special challenge to the TASI taxonomy. If there is no theory, then there are no theoretical assumptions with nonobservational terms, and there is no need for auxiliary assumptions to link those nonexistent nonobservational terms to observational terms in empirical hypotheses.
There are two responses to the challenge. First, it is possible to argue that there is always a theory, even if implicit, and so the TASI taxonomy applies even for ostensibly atheoretical research. 6 A second response could be based on accepting the existence of research sans theory. In this latter case, consider that although auxiliary assumptions will not fulfill one of their functions described earlier when there is no theory, which is to connect nonobservational terms in theories with observational terms in empirical hypotheses, they nevertheless can fulfill other functions, such as providing initial conditions and addressing confounding issues. For instance, imagine an applied experiment in which participants are provided with a pro-seatbelt essay in the experimental condition but not in the control condition, with a subsequent measure of seatbelt use. Although a theoretically inclined researcher might hope that the essay increases attitudes toward using seatbelts, thereby increasing seatbelt use, such an assumption is not required for applied research. From an applied perspective, it might be sufficient that the essay works, regardless of theory. Of course, it remains necessary to assume that the participants in each condition received the right forms, that they understood the directions, and so on. Although auxiliary assumptions cannot fulfill the connective function when there is no theory, they nevertheless fulfill other important functions and are ignored at the researcher’s peril.
Little more needs to be said about statistical and inferential assumptions, which fulfill much the same roles in applied research as they do in basic research. This is because, again, the empirical hypothesis likely is not sufficiently well specified, thereby necessitating statistical assumptions to facilitate a statistical hypothesis (e.g., the means in the two conditions differ). And there is still the issue of using sample data to draw conclusions about populations.
In terms of subscripts, taking out the theoretical assumptions still leaves auxiliary, statistical, and inferential assumptions. Instead of a modelTA, modelTAS, or modelTASI, we instead have a modelA, modelAS, or modelASI, respectively. Therefore, for non-theory-based applied research, the taxonomy simplifies by the removal of the first of the four subscripts and remains relevant. However, the issues of internal and external validity take on somewhat different characteristics. In theory-testing research, Sense 2 internal validity is about being able to attribute the effect to the manipulation having influenced the theoretically relevant construct. If the essay versus no-essay manipulation, for example, could reasonably be argued to manipulate something other than attitude that, in turn, influences behavioral intentions, then this is a serious Sense 2 internal-validity problem even when there is consensus that the essay manipulation causes the effect. It is not sufficient that the manipulation works (Sense 1 internal validity); it must work for the right reason (Sense 2 internal validity). In contrast, for atheoretical research, there is no theory by definition, and it is debatable whether the manipulation must work for the right reason to be counted internally valid. If the essay manipulation increases seatbelt use, the fact that it works might be sufficient, even if it is not clear why it works. In that case, one might reasonably deem the experiment to have good Sense 1 internal validity, and Sense 2 internal validity is irrelevant because of the lack of a theory.
The nature of external validity also changes in applied contexts sans theory. When applying external validity to theories, the idea is that the theory is useful in different cultures, contexts, research paradigms, and so on. But to make the theory generalize thusly, it might be necessary to vary the auxiliary assumptions, as described earlier. However, applied work, if not theory-based, may be different. In the case in which an essay increases seatbelt use, but for unknown reasons, and so the researcher hopes the essay will work again in another country, there is no theory to generalize. The researcher is forced into the position of attempting to generalize the finding, that is, Sense 1 external validity, and the difficulties detailed by Yarkoni (2022a) and some of the commentators retain their full force. There can be no Sense 2 external validity because there is no theory to generalize. This need not be an argument against applied research not based on theory; Lesko et al. (2020) provided an excellent description of the relevance of internal- and external-validity concerns for policy decisions, despite being limited to a Sense 1 internal- and external-validity perspective.
In summary, applied research may or may not be theory-based. When applied research is theory-based, the conclusion that better assumptions across the taxonomy benefit both Sense 2 internal and external validity still applies. When applied research is not theory-based, Sense 2 validity is out of the question and only Sense 1 validity remains. In that case the received view applies, and the researcher may need to perform a variety of different studies, some to maximize Sense 1 internal validity and some to maximize Sense 1 external validity, in the traditional spirit of Brinberg and McGrath (1985).
Typical Research Strategies to Follow (or Not)
We have already seen some new and important points. One is that although Sense 1 internal and external validity may often be in trade-off, consistent with the received view, Sense 2 internal and external validity practically always operate in unison. Second, when researchers argue about internal and external validity, it is crucial to distinguish whether the argument is along Sense 1 or Sense 2 validity grounds. Arguments that seem sound from one of the perspectives may not be sound from the other, and vice versa. For example, we saw earlier that Yarkoni’s (2002a) argument worked better from a Sense 1 external-validity perspective than a Sense 2 external-validity perspective, whereas his critics’ arguments worked better from a Sense 2 external-validity perspective than from a Sense 1 external-validity perspective. Recognizing this distinction suggests that multiple—and even seemingly contradictory—perspectives can have worth. The integration of the distinction between Sense 1 and Sense 2 validities with the TASI taxonomy facilitates the conversation about internal and external validity. Third, although it is a truism that internal validity is a prerequisite for external validity, there are times when the reverse is true. The ability to say something about populations is often a necessary condition for making a convincing case that the effect is due to the manipulation (Sense 1 internal validity) or for the correct theoretical reason (Sense 2 external validity). In short, external validity can be a prerequisite for internal validity.
Fourth, we have seen that the relationship between Sense 1 and Sense 2 internal validity differs from the relationship between Sense 1 and Sense 2 external validity. Sense 1 internal validity is a necessary but not sufficient condition for Sense 2 internal validity. In contrast, Sense 1 external validity is not even a necessary condition for Sense 2 external validity.
Fifth, there is a tendency for researchers to emphasize replication failures. From a Sense 1 external-validity perspective, this is understandable. However, from a Sense 2 external-validity perspective, we have seen how changing the auxiliary assumptions to be more context-appropriate can transform Sense 1 external-validity problems into Sense 2 external-validity virtues. If replication failures are found to be due to context-inappropriate auxiliary assumptions, and then replication successes ensue based on context-appropriate auxiliary assumptions, Sense 2 external validity can be greatly enhanced. Thus, a Sense 1 external-validity problem need not be negative; it can be positive if followed up by careful research with a strong focus on context-appropriate auxiliary assumptions.
These are conceptual gains. Are there direct implications for the conduction of research? A few are described in the following subsections.
Different sets of studies for Sense 1 internal and external validity (or not)
As a result of the received view of an internal–external validity trade-off, the traditional recommendation and strategy is for researchers to consider them separately and perform two sets of studies, one set to maximize internal validity and another set to maximize external validity. However, the present perspective suggests this advice involves a crucial but unrecognized cost. Put simply, even if we accept that the internal-validity studies really do establish Sense 1 internal validity and that the external-validity studies really do establish Sense 1 external validity, none of the single studies make a good case for both internal and external validity; this is the basic essence of the received view. Far worse, however, from a Sense 2 validity perspective, it is not clear why even the combination of studies makes a strong case for the theory. Accepting Sense 1 internal and external validity nevertheless leaves open that the studies may have worked for theoretically wrong reasons or that generalization occurs for theoretically wrong reasons. Thus, in terms of the goal, which is to make the most definitive statement possible with respect to the theory, the traditional strategy is wanting. This is not to say that Sense 1 validity is irrelevant. It is not irrelevant because if a researcher is not confident that an effect is due to the manipulation, that researcher likely will not be confident that the effect is due to the correct theoretical reason either. However, Sense 1 internal and external validity are far from sufficient.
An explicit acknowledgment that Sense 2 internal and external validity are crucial places the received view in its proper place. It is not that the received view is incorrect, only that its applicability is limited to Sense 1 validity. When we move to Sense 2 validity, with its emphasis on theory, we see that internal and external validity increase or decrease in unison, depending on whether the researcher makes high- or low-quality assumptions at relevant levels of the TASI taxonomy. This change in viewpoint implies the futility of separate internal- and external-validity studies and suggests researchers should consider Sense 2 internal and external validity simultaneously to make the most of their propensities for mutual buttressing. In turn, for this strategy to work, researchers must focus on how best to negotiate relevant levels of the TASI taxonomy to maximize Sense 2 internal and external validity.
To prevent misunderstandings, the recommended change in strategy is not an argument against multiple studies but rather an argument against separating the studies into those designed to maximize internal validity, mostly ignoring external validity, or to maximize external validity, mostly ignoring internal validity. Once we take a Sense 2 validity perspective, each study should focus on both internal and external validity simultaneously, with multiple studies providing opportunities to explore different ways to negotiate levels of the TASI taxonomy. Naturally, although perhaps not obviously, following this advice necessitates researchers carefully spell out as many of the theoretical, auxiliary, statistical, and inferential assumptions as feasible, not just for a single study but for all studies bearing on the theory of concern. It is worth reiterating that this is not merely a matter of “operationalizing” constructs but of consciously spelling out the auxiliary, statistical, or inferential assumptions that span the distances between levels of research. It might or might not be that researchers believe they consider the different levels of assumptions. However, any reasonable reading of the psychology literature renders indisputable the fact that researchers fail to do a good job of stating them explicitly. Without explicit statements of assumptions, evaluating their soundness is difficult. Thus, the point is twofold. First, researchers should eschew separate internal- and external-validity studies in favor of considering both in each study. Second, researchers should state not only the theory as clearly as possible but also the auxiliary, statistical, and inferential assumptions. Such clear stating of assumptions facilitates systematically changing them across studies to provide converging evidence if all are successful or to show that some types of assumptions are problematic, with the possibility of rethinking one’s commitment to the theory.
Correlational studies, with mediation, for external validity (or not)
As explained earlier, with respect to causation, mediation analyses are far from being as definitive as researchers typically take them to be (Fiedler et al., 2011; Grice et al., 2015; Kline, 2015; MacKinnon et al., 2000; Tate, 2015; Thoemmes, 2015; Trafimow, 2015). Nevertheless, it remains possible to point to an external-validity advantage. Although mediation is sometimes used to augment experimental research, most mediation studies are purely correlational; there is no laboratory manipulation, and so there is less of an issue of abstracting the variables out of their social contexts. Thus, Sense 1 external validity increases, which could be argued positive.
However, from a Sense 2 perspective, this external-validity advantage may disappear. One problem is in identifying which assumptions are at the theoretical level and which are at other levels of the taxonomy. Consider a typical theory of reasoned-action study, in which many variables are measured, among them behavioral beliefs (and accompanying evaluations), attitudes, and behavioral intentions. The mediation diagram would, among other arrows involving other variables, include an arrow from beliefs (and accompanying evaluations) to attitudes and from attitudes to behavioral intentions. Because beliefs (and accompanying evaluations), attitudes, and behavioral intentions are in the theory, the obvious conclusion is that the assumptions illustrated by the arrows are at the theoretical level of the TASI taxonomy. However, appearances may be deceiving. Consider that the theory of reasoned action does not insist that attitudes cause behavioral intentions because subjective norms can cause them as well, and descendants of the theory include yet more variables that can cause behavioral intentions. Let us say that the arrow from attitudes to behavioral intentions has a path coefficient of 0; this can be interpreted as consistent with the theory if one of the many other variables measured in the study, such as subjective norms, predicts behavioral intentions. Or let us say that the arrow from attitudes to behavioral intentions has a path coefficient very different from 0; this, too, can be interpreted as consistent with the theory. Thus, almost no matter how the path coefficients come out, with the obvious exception being that all arrows leading to behavioral intentions have path coefficients equal to 0, the data can be, and generally are, interpreted as consistent with the theory.
The internal-validity issues are too obvious to require discussion, and so the present point is that there are also Sense 2 external-validity problems that may be papered over by traditional Sense 1 external-validity thinking. As we just saw, one tricky issue is that a theoretical arrow need not imply the existence of an arrow in the study at hand. One would have to make auxiliary assumptions to decide whether a theoretical arrow should or should not manifest in the context of the type of behavior, population of concern, and so on with respect to a single study. It is not difficult to imagine that under some types of behaviors, populations of concern, and so on, the theoretical arrow should manifest whereas under other variants it should not. Researchers seldom consider the auxiliary assumptions necessary to make predictions about these variants; rather, if any arrow is statistically significant, the data are interpreted as supporting the theory. From a Sense 2 external-validity perspective, this is poor, though typical, practice and should be changed. Moreover, Sense 2 external validity could be enhanced if researchers were to be explicit about auxiliary assumptions implied by different variants, so that a pattern of predictions about when theoretical arrows should manifest or not could be tested empirically. Empirical failures, if there are any, could be investigated to determine whether the problem is at the level of auxiliary assumptions or at the level of theory.
To prevent misunderstandings, the point is not that researchers should avoid correlational research or mediation analysis. Rather, it is that there can be a Sense 2 external-validity problem if the researcher is not explicit about the auxiliary assumptions used to predict whether a theoretical arrow will or will not manifest in varying study contexts. Such specification is rare, but the present perspective renders clear that it ought to become regular practice for correlational research.
Other contexts confer external validity (or not)
Consider again the Festinger and Carlsmith (1959) cognitive-dissonance experiment. We saw earlier that although there is much agreement that the manipulation causes the effect (Sense 1 internal validity), there is substantial disagreement about whether this is because of cognitive dissonance (Festinger & Carlsmith, 1959) or self-perception (Bem, 1967, 1972), thereby exemplifying a Sense 2 internal-validity difficulty. But let us now imagine a hypothetical replication performed in another culture. If the replication fails, the failure need not militate against the theory if an auxiliary assumption is at fault, such as that a different amount of money is needed in the other culture to produce cognitive dissonance. If the replication succeeds when the proper amount of money for the culture is used, the Sense 1 external-validity failure can be transformed into a Sense 2 external-validity success. Thus, one piece of advice for researchers is not to take Sense 1 external-validity failures too much to heart because they can be converted to Sense 2 external-validity successes by improving the auxiliary, statistical, or inferential assumptions.
However, let us suppose that the cognitive-dissonance replication succeeds right away not only in the other culture but also 27 additional cultures, with no failures whatsoever. Such an astounding generalizability success rate, in additional cultures, would doubtless be taken to indicate strong external validity. But no matter how spectacular the string of successes, this only constitutes Sense 1 external validity. If the many successes are due to self-perceptions rather than cognitive dissonance, they would constitute Sense 2 external-validity failures. Thus, researchers should not take apparent external validity successes too seriously unless they have reason to believe that they are for the correct theoretical reason. This is not an argument for researchers to avoid different contexts or to avoid cross-cultural research. On the contrary, this is an argument in favor of that but with the crucial augmentation that auxiliary, statistical, or inferential assumptions have been carefully thought through, and perhaps tested empirically, to support Sense 2 external validity.
A compelling physics example pertains to Aristotle’s theory that objects fall because it is in their nature to fall, so heavier objects should fall faster than lighter objects. Observations of heavier objects falling faster than lighter objects in many contexts seemed to support external validity for almost two millennia until Galileo proposed better auxiliary assumptions pertaining to the interaction of the characteristics of falling objects and the Earth’s atmosphere. When Galileo used these better auxiliary assumptions with an ingenious experimental setup that rendered the interaction trivial, typical observations no longer replicated, thereby constituting a Sense 2 external-validity problem, although Sense 1 external validity remained. That is, outside Galileo’s highly artificial setup, heavy objects continued to fall faster than lighter ones, just as Aristotle predicted and the ancient Greeks consistently observed. In turn, Galileo’s Sense 2 external-validity problem—a failure to replicate using his artificial experimental setup—also caused a Sense 2 internal-validity problem, as it became clear that confirmations of Aristotle’s prediction in most normal life situations was not because of Aristotle’s theory. The historical fact is that Galileo’s inertial theory prevailed over Aristotle’s theory because physics researchers realized that Sense 2 external validity trumps Sense 1 external validity, which is prevalent in psychology. Physics researchers also realized that a Sense 2 external-validity failure is also problematic for Sense 2 internal validity. Psychology researchers would do well to internalize the Galilean lesson. Not all psychology difficulties are caused by difficult subject matter; some are caused by a widespread and insufficient understanding of internal and external validity that the present exposition hopefully helps to remedy.
Typical effect sizes (or not)
Let us return to the theory of prejudice introduced by Stephan and Stephan (2000) and a hypothetical experiment in which participants are exposed to an out-group-relevant threat (experimental condition) or not (control condition), with a subsequent measure of prejudice against that out-group from +3 (extreme prejudice) to −3 (extreme favorable feelings). Suppose that mean prejudice is +2 in the experimental condition and +1 in the control condition, and the standard deviation is 2 in both conditions. In this case, the standard effect-size measure, Cohen’s d (the difference in means divided by the standard deviation), would be as follows:
Despite our extreme generosity in making amazingly favorable assumptions, the extent of the Sense 2 internal and external validity is not necessarily clear because, among other issues, there was no mention of the shapes of the distributions or of the skewness of the data in each condition. Suppose that the data in both conditions come not from normal distributions but from skew-normal distributions and that the skewness in the experimental condition is 0.10 and −0.10 in the control condition. These skewness values are so close to 0 that most researchers would characterize the data as approximately normal in both conditions if they bothered to look at skewness at all. However, running out the calculations for estimates of skew-normal parameters rather than normal parameters gives the following estimated locations: experimental condition (0.77) and control condition (2.23). The estimated scale is 2.35 in both conditions, and the estimated shapes are as follows: experimental condition (0.87) and control condition (−0.87). The skew-normal effect size (the difference in locations divided by the scale) is −0.622. In summary, although the difference in means, along with the effect-size calculation, supports the theory, the difference in locations, along with the distribution-appropriate effect-size calculation, contradicts the theory.
Thus, despite what may have seemed airtight support for the theory under our too generous assumptions, we are nonetheless in a dilemma. The theory requires the experimental-condition distribution to shift upward relative to the control-condition distribution. In contrast, the difference in locations shows the opposite—a downward shift relative to the control condition. Worse yet, the distribution-appropriate effect-size calculation is also in the wrong direction.
In terms of Sense 1 internal validity, there can be no doubt that the experimental manipulation caused the difference in means, and so all appears well. And this would be true of Sense 1 external validity too were the difference in means to replicate in other countries, with other out-groups, and so on. But the experimental manipulation causes a difference in means for the wrong reason, and so all is not well. The experimental manipulation both introduces slight positive skewness and a location shift in the downward (theoretically wrong) direction. The statistical assumption of normality that underpins the difference in means and its associated Cohen’s d is false, there is a lack of generalization across statistical assumptions, and the prediction is disconfirmed under the appropriate skew-normal statistical assumption. Thus, the effect of the manipulation on the difference in means and its associated Cohen’s d is not for the theoretically correct reason. This issue of the difference in means and difference in locations being in opposite directions may seem fanciful, but my colleagues and I are collecting data sets, and based on the 50 we have collected thus far, effects in opposite directions have occurred in over a third of them.
There is an easy way to avoid such problems. Specifically, whenever researchers assume normality, they should also try out an assumption of skew normality. If normality is true, nothing is lost because normal distributions are a special case of skew-normal distributions. In that case skew-normal calculations will approximately equal normal calculations. However, it may well be that skew-normal calculations come out quite differently from normal calculations, in which case there is the issue of which calculations to emphasize. For theory-based research, such as the example involving Stephan and Stephan’s (2000) theory of prejudice, the skew-normal calculations are more apt. However, there are other contexts, such as applied ones, in which a distribution shift may not be required. In this case, means might (or might not) be quite appropriate. 7
Conclusion
The present goal was to propose an integration involving the two senses of internal validity and external validity and the TASI taxonomy. As we saw earlier, these are inextricably linked: There is no Sense 2 internal or external validity without negotiating levels of the TASI taxonomy, and the relevance of the TASI taxonomy is augmented by a concern with Sense 2 internal and external validity. This integration helps address admonitions, such as Yarkoni’s (2022b), that calls for more theory are unhelpful. We now have a clear specification of the different levels of assumptions needed to traverse the distance between theories and empirical, statistical, or inferential hypotheses to maximize Sense 2 internal and external validity. An underestimated difficulty not only in psychology but also in other social sciences is that most assumptions across the levels of the TASI taxonomy are unconscious, unstated, or both. An important consequence is that such assumptions are difficult to investigate. The present integration hopefully will stimulate researchers to think consciously about their assumptions across the TASI taxonomy, and state them too, thereby bringing them out into the open for thorough examination.
Of course, there is always the issue of how one arrives at a theory, to begin with, which requires its own space to address properly. However, the present integration suggests an intriguing possibility, which is to mentally toggle across levels of the TASI taxonomy. Although it was convenient to start with theory in the foregoing explanation of TASI, it is quite reasonable to start at other levels. For example, when reading an article, a researcher may look for, and question, the auxiliary, statistical, or inferential assumptions. In turn, such questioning can lead to doubts about Sense 2 internal or external validity, thereby activating potential alternative assumptions and perhaps even stimulating the proposal of an alternative theory. By continually mentally toggling across the different levels of assumptions, researchers can make improvements at multiple levels, including the theoretical level. Thus, the present perspective is optimistic. A focus on the different levels of assumptions really can substantially improve psychological research, provided researchers (a) state the assumptions being made at each level so they can be meticulously examined and (b) search for new assumptions that could be brought to bear on an issue, to increase the auspiciousness of constellations of assumptions across levels of the TASI taxonomy in future research.
Footnotes
Acknowledgements
I thank John Richters for his continual interest, a set of stimulating phone discussions, and high-quality help. I also thank Klaus Fiedler for providing an unusual and valuable level of editor engagement.
Transparency
Action Editor: Klaus Fiedler
Editor: Klaus Fiedler
