Abstract
There is no direct way for researchers to test theories. One reason is that theories contain nonobservational terms that refer to unobservable entities. Consequently, researchers add auxiliary assumptions to aid in traversing the distance between theories and empirical hypotheses. The results may confirm the empirical hypothesis, an empirical victory, or the results may disconfirm the empirical hypothesis, an empirical defeat. Either way, it is not clear whether to make an attribution to the theory, the auxiliary assumptions, or both. The present goal is to review techniques researchers have employed, or could employ, that aid in assessing the weight of the evidence with respect to crediting or blaming theories or auxiliary assumptions for empirical victories or defeats.
As the philosopher Butterfield (1957) described, there has long been controversy about accepting versus rejecting the implications of experimental findings, and such controversy was particularly marked during the 17th century. Even Galileo (1564–1642), for all his vaunted empiricism, admitted to having several times dropped lumps of lead and wood from a tower in his youth, with the result being that a lump of lead always leaves the lump of wood behind. The replicable finding clearly supports Aristotelian (384–322 BC) theorizing that objects fall because it is in their nature to fall, and heavier objects (e.g., a lump of lead) have more of this nature than do lighter objects (e.g., a lump of wood). Nevertheless, Galileo did not allow the finding to prevent him from eventually theorizing to the contrary and devising better experiments with contrary findings. Butterfield’s rendition of Galileo reveals a tension. On the one hand, it is good to provide empirical tests of theories. On the other hand, sometimes it is good not to be too influenced by empirical findings.
Modern philosophers and philosophically inclined psychologists agree with the increased emphasis in the 17th century on empirical findings. Today, the desirability of testing theoretical conjectures against empirical facts is a truism, though another truism is that there are few crucial studies where the results unambiguously confirm or disconfirm the theories they purport to test (Duhem, 1914/1954; Hempel, 1945). A source of ambiguity is that theories contain nonobservational terms referring to unobservable entities (Cronbach & Meehl, 1955; Lakatos, 1978; Meehl, 1990a; St Quinton et al., 2021; Trafimow, 2009, 2019b). In contrast, empirical hypotheses contain observational terms referring to observable entities. A well-designed experiment might provide convincing evidence with respect to causal relations among observable entities in an empirical hypothesis, but there remains distance to be traversed to draw inferences about causal relations between unobservable theoretical entities. For a quick example, to test whether attitudes cause intentions to perform behaviors (e.g., Fishbein & Ajzen, 1975, 2010), it is possible to randomly assign participants to receive a pro attitude essay or a neutral essay, and then measure intentions to perform the focal behavior. Even if the experiment works as planned, so that intentions are more positive in the pro than neutral condition, there is no guarantee that the essay manipulation influences attitudes and not something else (e.g., mood), that the intention measure measures intentions and not something else (e.g., social desirability), and so on. These additional assumptions, linking the observable manipulation to the unobservable attitude, linking the observable intention measure to unobservable intentions, and so on, are called auxiliary assumptions. Additional auxiliary assumptions are necessary to set initial conditions, such as that the research assistant passed the correct forms to the correct participants, that the process of random assignment really was random, and so on. In the end, if the study fails, it is difficult to know whether to attribute the empirical defeat to the theory being false or to having at least one false auxiliary assumption. Or, if the study succeeds, it is difficult to know whether to attribute the empirical victory to the theory being true or having at least one wrong auxiliary assumption (e.g., the manipulation influenced mood, rather than attitude, which caused the effect on the intention measure). Thus, auxiliary assumptions are important to appraise the validity of both empirical failures and empirical victories.
This quandary is the present topic. It would be desirable to present a clear decision rule that unambiguously states, in the event of an empirical victory, whether the victory was obtained because of a correct theory and correct auxiliary assumptions or whether the theory is false but an incorrect or extremely powerful auxiliary assumption caused that victory anyhow. For instance, in the attitude example, do attitudes really cause intentions or is the observed effect for some other reason?
Or suppose that the attitude study fails. Is it because attitudes do not cause intentions or because the attitude manipulation fails to manipulate attitudes, the intention measure fails to measure intentions, the research assistant passed out the forms incorrectly, and so on? In the event of an empirical victory, can we soundly credit the theory? In the event of an empirical defeat, can we soundly blame the theory?
Unfortunately, we know of no single and unambiguous decision rule that provides answers to these questions. However, that is no reason to give up and, given the importance of the issue, nor should we. There are a variety of techniques that can provide relevant evidence, even if the evidence is not 100% definitive. The present goal is to review some techniques and discuss them. However, we first acknowledge the potential relevance of what might be considered a preliminary issue concerning how researchers arrive at empirical hypotheses and the attendant difficulties. A second preliminary issue concerns the possibility that auxiliary assumptions can be made part of the theory, thereby rendering difficult discerning central versus auxiliary assumptions.
Arriving at empirical hypotheses
There is an indefinitely large set of potential auxiliary assumptions to connect with a theory, which creates two immediate problems. One problem is generating potential auxiliary assumptions; failure to generate high-quality ones is deleterious to the research process. Secondly, even if a researcher generates a list of auxiliary assumptions that includes high-quality ones, there may be a failure to select the better ones out of the list. Thus, there is a generation problem and a selection problem.
An additional difficulty is that a theory may imply multiple empirical hypotheses, depending on which auxiliary assumptions researchers employ (Katzko, 2006). Too, a confirmed empirical hypothesis can be implied by multiple, and perhaps even contradictory, theories, again depending on the auxiliary assumptions that researchers employ (Spirtes et al., 2000). The difficulties in both directions can prove challenging when researchers wish to derive empirical hypotheses from theories or when they wish to induce theories from findings in psychological literatures.
Katzko (2006) provided many examples pertaining to automaticity, including the famous finding that priming the elderly decreases walking speeds for participants upon leaving the laboratory (Bargh et al., 1996, Experiment 2). Leaving aside the issue of the replicability of the finding (e.g., Doyen et al., 2012), there is the logical issue of including sufficient auxiliary assumptions to entail the prediction from the theory. The theory, in brief, is that the presentation of a stimulus is sufficient to “automatically” induce behavior. In this particular instantiation, the empirical hypothesis is that priming the elderly should reduce walking speeds. However, as Katzko indicated, it is necessary to assume that slow walking is part of the stereotype of the elderly. And even if this is granted, another necessary assumption is that this particular aspect of the stereotype is active as opposed to some other aspect of the stereotype. Stated more generally, insisting that people act automatically in accordance with the stereotype fails to specify exactly how participants are supposed to act and, perhaps as important, how participants are not supposed to act. Which components of the stereotype should or should not result in automatic behavior, and under what conditions? Even granting the finding, the reasoning gaps Katzko identified leave quite debatable the extent to which that finding provides convincing support for Bargh et al.’s automaticity assertion.
The difficulty in moving from theories to empirical hypotheses designed to test them was explained in another way by Kellen (2019; also see Kellen et al., 2021), who focused on different sorts of models that are useful for bridging the distance between theories and empirical hypotheses. Models encompass sets of auxiliary assumptions that might or might not be correct. Although Kellen (2019) identified several kinds of models, measurement models provide an illustrative case. A typical assumption that measurement models in the personality area employ, though usually implicit, is transitivity (also mentioned by Kellen, 2019 in a different context). For example, if A is more neurotic than B, and B is more neurotic than C, then A is more neurotic than C. However, Morris et al. (2017) tested the transitivity assumption with respect to neuroticism and found an impressive proportion of transitivity violations, thereby calling the transitivity assumption into question. The point is not that transitivity works or does not work in personality psychology, but that it is one of many assumptions in typical measurement models that might be false or at least not always true.
The potential incorporation of an auxiliary assumption into the theory
It sometimes is not clear when an auxiliary assumption becomes a central component of the theory. For example, consider terror management theory (TMT), according to which people have a fear of death that can create paralyzing terror that they deal with in a variety of ways, such as developing and maintaining a cultural world view and self-esteem (Becker, 1973; J. Greenberg et al., 1990). Many researchers have obtained what might be termed TMT effects by randomly assigning participants to contemplate their eventual death (experimental condition) or dental pain (control condition). Typically, participants reinforce their own world-view, denigrate outgroups, and a variety of others, more in the mortality salience than dental pain condition. However, an important problem is that most TMT effects only occur when there is a delay between mortality salience and the main dependent variable. TMT proper does not predict that TMT effects should depend on a delay, and so TMT theorists developed the death-thought-suppression-and-rebound assumption. The idea is that immediately after mortality is made salient, participants suppress death thoughts, but they do not stay suppressed, and inevitably rebound after a delay (e.g., T. Greenberg et al., 1994; J. Greenberg et al., 2000; Pyszczynski et al., 1999). When the death-thought-suppression-and-rebound assumption is added to TMT, then findings of TMT effects only after a delay, but not immediately, can be interpreted as strongly supporting rather than inconveniencing the theory. Question: should the death-thought-suppression-and-rebound assumption be considered central or auxiliary to the theory? If the assumption is considered central, then disconfirming it would be strongly problematic for the theory. In contrast, if the assumption is considered auxiliary, then disconfirming it would only be mildly problematic for the theory; it might be reasonable to find another auxiliary assumption, or set of auxiliary assumptions, with which to replace it and thereby save the theory.
The issue is not merely academic. Trafimow and Hughes (2012) performed a set of experiments explicitly designed to test whether death thoughts following mortality salience treatments are more accessible immediately or following a delay. In direct opposition to the death-thought-suppression-and-rebound assumption, these researchers found that death thoughts were more accessible immediately than after a delay. Whether the findings from these experiments should be considered as strongly or weakly problematic for TMT depends importantly on whether the death-thought-suppression-and-rebound assumption is considered central or auxiliary. However, we hasten to remind the reader that even if the death-thought-suppression-and-rebound assumption is deemed auxiliary, the evidence against it begs for an alternative explanation of the findings that had led to it, and the present authors are unaware of any. As a matter of historical fact, most TMT researchers have simply ignored that the death-though-suppression-and-rebound assumption has been convincingly disconfirmed, which illustrates a sociology of science problem that psychology researchers enamored with a theory may ignore problematic findings. 1
It has happened, though rarely, that theorists have unambiguously incorporated what otherwise would be auxiliary assumptions into the theory. Fishbein and Ajzen (Ajzen, 1988; Ajzen & Fishbein, 1980; Fishbein, 1980; Fishbein & Ajzen, 1975) provided a famous example, by including the principle of correspondence or compatibility into the theory of reasoned action. This is a measurement principle, according to which all constructs in the theory should be measured at the same level of specificity. 2 Indeed, Davidson and Jaccard (1975) showed that the theory’s predictive power is greatly enhanced when the constructs are measured in accord with the principle.
For a thermometry example, Kellen et al. (2021) considered the development of thermometers and the difficulty in connecting the issue of what temperature is and how to measure it. After centuries of work, the solution eventually came in the form of a theoretical advance, the kinetic theory of heat, which redefines temperature as the average kinetic energy of the particles in an ideal gas. From one perspective, the kinetic theory of heat can be considered an auxiliary assumption, though an important one, adding to the body of work on thermometry. However, from another perspective, the kinetic theory of heat arguably became much more than simply an auxiliary assumption to aid in the construction of thermometers; the kinetic theory was crucially important in the history of 19th-century science and arguably should be considered central.
In summary, auxiliary assumptions may become a sufficiently important part of scientific thinking to become central and no longer auxiliary. That which counts as central or auxiliary, though sometimes obvious, need not always be so.
When there is an empirical victory
Suppose a researcher tests a theory and the findings support that theory; there is an empirical victory. It would be nice to attribute the empirical victory to the theory and auxiliary assumptions being true, but even false theories or auxiliary assumptions can result in empirical victories. This major section concerns techniques to help make the case for crediting the theory.
A few traditional techniques
In this subsection, we discuss a few traditional techniques. The idea is not to cover the field—we leave that to methods textbooks—but rather to provide an intellectual flavor that hopefully will carry over to techniques not discussed, too.
Random assignment of participants to conditions
Methods textbooks emphasize random assignment of participants to conditions (MacLin, 2020). The hope is that such random assignment renders the participants in the different conditions equivalent (or approximately equivalent) on all causally relevant variables. Thus, if a critic wishes to attribute an empirical success to a lucky distribution of participants across conditions, the retort is that random assignment renders this implausible. With the alternative explanation rendered implausible, the case for crediting the theory improves.
It is possible to overstate the case for the efficacy of random assignment of participants to conditions. One issue is that there is usually no way to know the number of causally relevant variables and so it is far from clear whether random assignment really renders the desired initial equivalence across groups. Another issue is the generalization issue. Although random assignment of participants to conditions can be argued to justify generalization to the population of randomizations that could have been performed, this is very different from generalizing to populations of people. To generalize to populations of people, it is necessary to randomly select from them, and this is rarely (or perhaps never) accomplished in psychology. In turn, if there is no sound way to generalize to a population of people, then theories that make assertions about populations of people can be argued to be poorly tested. The empirical success might better be credited to idiosyncrasies of the sample than to the theory. Despite the caveats, random assignment of participants to conditions is generally desirable, when feasible.
Manipulation checks
Possibly the most obvious way to check on an auxiliary assumption pertaining to the manipulation is to perform a manipulation check (Fiedler et al., 2021). To recycle the attitude example, in addition to having a pro or neutral essay, and measuring intentions to perform the behavior, the researcher can add a measure of attitudes. If the manipulation really influences attitudes, the expectation would be that attitudes should be more positive in the pro condition than in the neutral condition. Suppose that this happens, both attitudes and intentions are more positive in the pro than neutral condition. How strongly have we supported the auxiliary assumption that the manipulation really influences attitudes, and nothing else of causal relevance?
If we assume that the attitude measure is valid (a potentially questionable assumption, but let us make it anyhow), then the effect of the manipulation on this measure can be considered as strongly supporting that the manipulation influences attitudes. That is the good news. The bad news is that the hoped-for effect on the manipulation check does nothing to reduce the plausibility that the manipulation influenced some other causally relevant variable too, such as mood. 3 Suppose that the manipulation influences both attitudes and moods. In that case, the effect of the manipulation on intentions could be for the theoretically correct reason, attitudes, or for a theoretically incorrect reason, moods. Thus, the manipulation check may not be as convincing as initial appearances suggest.
A way around this problem is to add a mood measure and hope for only a trivial effect of the manipulation on that measure. Under the perhaps questionable assumption that the mood measure is valid, a trivial effect on the mood measure coupled with an impressive effect on the intention measure would render mood implausible as an alternative explanation.
Another way to think about the issue might be based on Meehl’s (1990b) assertion that a large set of auxiliary assumptions, rather than directly implying an observation, implies a material conditional; based on the theory and auxiliary assumptions, if one observation is made, then another observation should be made. In that case, a potential interpretation of a manipulation check failure is that the antecedent of the material conditional has not been satisfied, thereby militating against using the material conditional to draw conclusions about the conjunction of theory and auxiliary assumptions. That said, it is possible to argue that the material conditional is inappropriate. An alternative might be to treat an experiment as employing a subjunctive conditional (based on the theory and auxiliary assumptions; if one observation were to happen, then another observation should happen). A potential advantage of the subjunctive conditional over the material conditional is that it could be argued to accord better with natural language (see Sanford, 1989 for a review). A disadvantage of the subjunctive conditional is that it is not subject to a truth table, thereby rendering vaguer the logic of its application to drawing conclusions about theories. Moreover, even aside from pitting the material conditional against the subjunctive conditional, it is far from clear that Meehl’s assertion of a theory and set of auxiliary assumptions setting up the material conditional is the best way to think about how researchers move from theory to empirical hypotheses (see discussion by Rozeboom, 2005). There will be no attempt here to argue for or against Meehl’s use of the material conditional, the subjunctive conditional, or others. It is sufficient that even the combination of an empirical victory with respect to the main dependent variable and the manipulation passing a manipulation check need not impressively enhance support for the theory.
Replication
With the current and widespread concern about whether studies replicate, the need for replication studies seems obvious (Nosek & Errington, 2020). We, too, believe that replication studies are desirable. However, they are far from foolproof. A study might replicate consistently for the wrong theoretical reason, in which case even many impressive and diverse replications may fail to adequately support the theory (Trafimow, 2019c). The most obvious example comes not from psychology but from the physics of falling objects. As we saw earlier, Aristotle proposed what might be considered a “trait” theory, which is that objects fall because it is in their nature to fall. The heavier the object, the more of this nature they have, and the faster they fall. Although the ancient Greeks tended to depend more on reasoning than on formal experimentation, it is easy to imagine experiments where they dropped iron skillets versus feathers, spears versus scraps of papyrus, and so on. And it is similarly easy to imagine dropping light versus heavy objects from many different heights, in many different places, with as many variations as desired. The extremely replicable finding would be that heavy objects really do fall faster than light objects, but not because of Aristotle’s theory. 4 Rather, the heavy objects’ speed of descent is less influenced by the Earth’s atmosphere. The problem is that if atmospheric effects are present in all the replications, as they would be other than in experiments such as Galileo’s later ones that were carefully designed to minimize these effects, then there would be an impressive degree of replication. We might be sorely tempted to wrongly conclude that Aristotle’s theory has been impressively supported and fail to understand the superiority of Galileo’s theory. Thus, although we support replications, that support has limits.
In addition to the potential problem that a false auxiliary assumption permeates many replications, there is a theoretical issue, too. Specifically, due to underdetermination with respect to auxiliary assumptions, it is more than possible for a false theory to result not only in replicable findings but in extremely important ones. In the history of chemistry, consider phlogiston theory, that combustion is due to the consumption of phlogiston. Under the aegis of phlogiston theory, many important and replicable empirical discoveries were made, including the discoveries of nitrogen and oxygen. Although Lavoisier (1743–1794) eventually disconfirmed phlogiston theory to researchers’ satisfaction, it is instructive to consider oxygen and nitrogen prior to Lavoisier’s devastating disconfirmation. At that time, oxygen and nitrogen were considered dephlogisticated and phlogisticated air, respectively. Objects were thought to burn well in what we now call oxygen because of the lack of phlogiston, and consequently the acceptance of phlogiston with unusual eagerness to facilitate combustion. In contrast, objects were thought to burn poorly in what we now call nitrogen because of phlogiston saturation, thereby rendering the air unsusceptible to accepting further phlogiston to retard combustion. In psychological contexts, Regenwetter and Robinson (2017) explicated reasoning fallacies with respect to behavioral decision research and Rotello et al. (2015) made the point with respect to multiple areas in psychology. Even important and replicable empirical discoveries are not immune to poor reasoning or poor auxiliary assumptions that persist across replications.
Sample sizes
There is a large literature admonishing researchers to collect sufficiently large sample sizes to have a good chance of obtaining statistically significant findings. And power analysis is recommended as the way to accomplish this (e.g., Cohen, 1988). However, there is a large literature criticizing significance testing. 5 If one rejects significance testing, power analysis becomes pointless. There is little reason to care about obtaining a sufficient sample size to have a good chance of obtaining statistically significant findings if one is not planning on performing a significance test.
To make the present point, it is not necessary to take sides in this debate. Even if one supports significance testing, there is a much better reason to care about sample sizes than to perform a significance test. To approach this reason, imagine that Laplace’s omniscient and truthful demon were to appear and proclaim that there is no relationship between sample statistics and corresponding population parameters. Thus, for example, sample means have nothing to do with corresponding population means, sample standard deviations have nothing to do with corresponding population standard deviations, and so on. In that case, no psychology findings would matter, as there would be no way to draw conclusions about anything beyond the experiment at hand. There would be no sound way to test theories because the results in a replication study would be as likely to turn out in the opposite direction as in the obtained direction. Thus, the demon’s proclamation illustrates, in dramatic fashion, the cruciality of faith that sample statistics provide reasonable estimates of corresponding population parameters.
There have been recent and promising developments, pertaining to the a priori procedure (APP), that focus on sample sizes not for power analysis, but rather for using sample statistics to estimate corresponding population parameters (see Trafimow, 2019a for a review). Depending on the assumptions the researcher wishes to make about population distributions (e.g., the population distribution is normal, lognormal, gamma, exponential, skew normal, etc.), there are APP equations available for computing sample sizes necessary to meet a priori specifications pertaining to confidence and precision. For example, suppose a researcher wishes to obtain a sample mean that has a 95% probability (confidence specification) of being within one-tenth of a standard deviation of the population mean (precision specification), and assumes the population is normally distributed. 6 In that case, an APP calculation would indicate that the researcher needs to collect 385 participants to meet both specifications. The good news is that if the researcher collects 385 participants, they can be confident that the sample mean will precisely estimate the population mean. Furthermore, once the data are collected, the researcher can check the assumptions and adjust the calculation accordingly.
But even the APP is not a panacea. If the theory or an auxiliary assumption is problematic, that the sample statistics accurately estimate corresponding population parameters may be irrelevant. To put this in the form of a pointed question: “What good is it to have accurate estimates of population parameters if those population parameters fail to provide a fair test of the theory?” Nevertheless, we support researchers using the APP to address the issue of sample sizes, provided they keep the pointed question in mind.
Mediation analyses
Although mediation analyses usually feature in correlational studies, they are fairly often used to support that an experimental manipulation works through the theoretically correct variable. In our attitude paradigm, the goal might be to have a path diagram where there is an arrow from the manipulation to attitude, and another arrow from attitude to intention, with hopefully a trivial “direct” arrow from the manipulation to intention. How well does the mediation analysis support that the effect of the manipulation on intention is for the theoretically correct reason? Well, not very well. If the effect is through mood rather than attitude, that would be consistent with the path structure. An exception would be if mood is measured with a trivial effect, as we saw earlier. Without a mood measure or manipulation, there is no way to perform a path analysis to convincingly eliminate mood as an alternative explanation.
Let us extend to cross-sectional correlational designs, the usual way mediation analyses are used. To adapt the attitude example, suppose the researcher measures attitudes towards a behavior, intentions to perform that behavior, and actual performance or nonperformance of the behavior. The hope is to establish a nice path structure where there is a nontrivial arrow from attitudes to intentions, and one from intentions to behaviors, but a trivial “direct” arrow from attitudes to behaviors. Suppose the hope is realized with a clear empirical victory. How convincingly does the empirical victory support the theorized chain of causation where attitudes cause intentions which, in turn, cause behaviors? The evidence for this indirect effect is extremely weak. It could be that intentions cause attitudes and behaviors, that intentions cause attitudes that cause behaviors, and countless other possibilities. Kline (2015, p. 205, Figure 2) showed that there are at least 18 plausible possibilities in a three-variable cross-sectional design (see also Tate, 2015; Thoemmes, 2015). This exemplifies the statistical indistinguishability problem; when multiple statistical models fit the data, it is difficult to know which to favor (Spirtes et al., 2000).
Of course, the researcher could test all 18 of Kline’s possibilities to see which work and which do not, but researchers rarely test alternative models. Worse yet, it could be that none of the possibilities are true. For instance, consider two blatantly incorrect models concerning the planets of the solar system. First, planetary velocity could cause planetary mass which, in turn, causes planetary momentum. Second, planetary mass could cause planetary velocity which, in turn, causes planetary momentum. Trafimow (2015) tested these false competing models and found strong support for the first false model but not the second false model. More generally, mediation analysis is an extremely weak form of evidence, but due to its convenience, most researchers do not comprehend the extent of the weakness. As the planets example illustrates, there are too many ways for correlated variables to work out in the hypothesized way even when the model is false.
To summarize, researchers routinely use mediation analyses. In experimental paradigms, a typical goal is to support that the manipulation works through the theoretically correct construct. Put in terms of an auxiliary assumption, the goal is to support that the manipulation manipulates the theoretical construct it is supposed to manipulate. However, we have seen that mediation analysis is too weak to make this argument definitively; there is no way to convincingly disconfirm alternative explanations without manipulating or measuring the constructs such alternative explanations feature. Nevertheless, though not particularly convincing, mediation analysis is at least more convincing in experimental than correlational research paradigms because of the initial plausibility that the manipulated variable is causal, even if it is not clear that such causation propagates through the touted path model. In correlational research paradigms, even this initial plausibility is compromised, thereby rendering many more potential path models plausible alternative explanations. Finally, whether the research paradigm is experimental or fully correlational, it is possible that both the independent variable and the outcome variable determine the ostensible mediating variable, a type of collider effect (e.g., Holmberg & Andersen, 2022; Pearl & Mackenzie, 2018). In that case, because the ostensible mediator is highly correlated with both the independent variable and the outcome variable, a typical path analysis likely will support the anticipated mediation even though there is no mediation. Mediation researchers should pay more attention to the possibility of collider effects, specifically, and alternative models, generally, than they do.
Rarely used techniques
There are rarely used, and in our opinion underused, techniques that can aid in determining whether to interpret empirical victories as being due to the theory or something else.
Holding the second variable constant
Consider any ABC type theory, where one variable (A) is theorized to cause another variable (B) which, in turn, is hypothesized to cause another variable (C). This is analogous to three dominoes, where the first domino tips the second domino which, in turn, tips the third domino (Grice et al., 2015). Continuing the domino analogy, suppose a person holds the second domino in place, so it cannot tip even upon being impacted by the tipping of the first domino. In that case, the third domino would not tip either. Returning to tests of ABC theories, although these are generally tested via mediation analyses, there is a more convincing way that is sometimes feasible. Suppose that the experimenter finds a way to hold the second variable constant, possibly by performing an experimental manipulation to drive it to a floor or ceiling. In that case, manipulating the first variable should have little effect on the third variable. If manipulating the first variable influences the third variable when the second variable is free to vary, whereas the effect of the first variable on the third variable is greatly attenuated when the second variable is much less free to vary, then the combination of strong effect and attenuated effect can provide powerful evidence for the ABC theory.
Trafimow et al. (2005) provided examples in the attribution area. Their ABC theory was that violations of Kant’s perfect duties cause more negative affect than violations of Kant’s imperfect duties. In turn, negative affect causes strong trait attributions. In one test, they performed an independent negative affect manipulation to drive negative affect to a ceiling or allow it to vary freely. When negative affect was allowed to vary freely, presenting participants with perfect or imperfect duty violations strongly influenced trait attributions. But when negative affect was driven to a ceiling, trait attributions were generally strong, and not dependent on type of duty violation.
Then, too, in another experiment, Trafimow et al. (2005) used a misattribution paradigm to induce participants to attribute negative affect to an irrelevant source or not. When there was no misattribution, so negative affect varied freely, there was a strong effect of violation type on trait attributions. However, in the misattribution condition, where “relevant” negative affect was driven to a floor, trait attributions were weak regardless of violation type.
Although Trafimow et al. (2005) performed traditional mediation analyses too, these authors indicated that those analyses provided only weak support for the theory. As we saw earlier, there are too many ways a false theory can nevertheless result in empirical victories with respect to mediation analysis. Hence, they used ceiling or floor effects to hold the second variable—negative affect—constant. This technique of allowing the second variable to vary freely, or not, provides a more definitive test of ABC theories than does traditional mediation analysis. A caveat is that the researcher must be sufficiently imaginative to find a way to drive the second variable to a ceiling or floor. Another caveat is the necessity to make another assumption, which is that there really is a floor or ceiling with respect to the construct of interest, and that the floor or ceiling is not an artifact of truncation with respect to the measuring device. In the case of negative affect, it is sensible that there is a floor; it seems possible, at any given time, for a person to experience zero or near-zero negative affect. Ceilings may be more complex. If there is no such thing as infinite negative affect, then negative affect has a ceiling as well as a floor. If one is willing to take the position that a person can experience infinite negative affect, then perhaps negative affect does not have a ceiling. Alternatively, negative affect might asymptote, so that what we are calling a ceiling really refers to reaching asymptote. And reaching asymptote may be sufficient to justify the recommended strategy. In general, researchers who wish to drive putative mediating variables to a floor or ceiling might profitably consider whether the construct has a floor or ceiling, or at least an asymptotic level of floor or ceiling.
Competing theories
A well-known way to support a theory is to test it against a competing theory (e.g., Platt’s method of strong inference, 1964). If the theory to be supported, in conjunction with auxiliary assumptions, predicts an effect that the competing theory does not predict, an empirical victory is considered to support the favored theory over the competing theory. For example, Stephan and Stephan (2000) theorized that feelings of threat cause prejudicial attitudes towards outgroups. Suppose a researcher were to manipulate threat and obtain an effect on prejudice towards an outgroup. 7 The researcher argues that the theory of reasoned action (Fishbein & Ajzen, 1975, 2010) predicts a null effect and so the findings support the threat theory over the theory of reasoned action. However, a problem with the example is the failure to distinguish between a theory predicting a null effect versus simply not making a prediction. The theory of reasoned action is about attitudes towards performing behaviors, not about attitudes towards groups of people, and so it does not predict a null effect; it does not make any prediction whatsoever. More generally, in many cases where a researcher tests allegedly competing theories, there is no competition whatsoever; the competing theory is a straw person theory that fails to make a prediction as opposed to making a false null prediction.
An underappreciated difficulty in using the strategy of competition is that different theories likely have different constructs. Different constructs necessitate different auxiliary assumptions to traverse the distance to observational empirical variables. Thus, the competition is not between two theories, but rather between one conjunction of theory and auxiliary assumptions versus another conjunction of a theory and different auxiliary assumptions. Is the empirical victory for the favored theory because it really is superior to the competing theory, or is the empirical victory due to connecting the competing theory with low-quality auxiliary assumptions?
There are ways to address this dilemma. If a researcher wishes to test a favored theory against a competing theory by means of disconfirming a null result, then that researcher should take great pains to ensure that the conjunction of competing theory and auxiliary assumptions really does predict a null effect, as opposed to simply not making a prediction. Alternatively, a creative researcher might be able to find auxiliary assumptions that, when combined with the competing theory, make a prediction in the opposite direction to the prediction emanating from the conjunction of the favored theory and auxiliary assumptions. Not only does this guard against mistaking a lack of a prediction for a null prediction, but it also adds a qualitative characteristic to the analysis in that the predictions differ in kind as well as in degree. An empirical victory, in this case, is even more impressive because it is less plausible to make an attribution to powerful versus less powerful auxiliary assumptions. However, even this strategy is not foolproof because the possibility always remains that one or more auxiliary assumptions connecting to the competing theory are wrong, thereby negating that the empirical victory eliminates the competing theory. Despite the complications, we believe that it is generally valuable to test competing theories.
Spelling out auxiliary assumptions
One way to help researchers decide whether to attribute empirical victories to the theory is to render the auxiliary assumptions as explicitly as feasible. If confronted with a set of explicitly stated auxiliary assumptions, a reviewer, editor, or reader can make a more informed judgment about their plausibility. In contrast, if the auxiliary assumptions are not spelled out, then it is difficult to assess their plausibility. Instead, one is faced with the daunting task of attempting to uncover unstated auxiliary assumptions to determine their plausibility.
Although we believe that psychology would benefit immensely from researchers spelling out their auxiliary assumptions a priori, there are limits, nonetheless. For example, it seems rather silly to explicitly state that one is assuming that the participants in the different conditions received the appropriate forms, that the data were entered correctly, and so on. In the realm of extremity, it is possible to argue that an auxiliary assumption is that there are no beings with magic powers causing the results. The larger point is that it is impossible to list all auxiliary assumptions needed to test the theory, and so judgment is needed about which ones to list or not list. Nevertheless, we believe that more in the direction of spelling out auxiliary assumptions would benefit psychology research.
When there is an empirical defeat
Suppose a researcher employs auxiliary assumptions to connect a theory to an empirical hypothesis, but the result is an empirical defeat. How can the researcher decide whether to blame the empirical defeat on the theory or on at least one low-quality auxiliary assumption?
It is a sociology of science truism that it is difficult to publish empirical defeats, especially if the empirical defeats are accompanied by a lack of statistical significance. This is not to say that empirical defeats are never published; they are. But publication frequencies are overbalanced, in the extreme, in the direction of publishing empirical victories (Fanelli, 2012; Franco et al., 2014; Joober et al., 2012). A consequence is that there is very little in the way of what might be considered standard ways to handle empirical defeats. And yet, most researchers will admit that many of their studies fail. Rather than throwing these in the circular file, perhaps they can be made valuable. If the researcher could gain evidence that the theory is to blame, that would provide a reason to propose a new theory, whereas blaming the auxiliary assumptions would provide a reason to invent better ones.
Manipulation checks, revisited
We saw earlier that manipulation checks are limited in their usefulness because even given that the manipulation influences the desired variable, a manipulation check, by itself, is insufficient to eliminate the possibility that the manipulation influences some other variable too which, in turn, influences the dependent variable. But this argument pertained to empirical victories. For empirical defeats, manipulation checks can be much more powerful.
While proceeding towards an empirical defeat with respect to the attitude experiment, suppose the researcher includes an attitude manipulation check and finds that the pro attitude essay has only a trivial effect on attitudes. Assuming that the attitude manipulation check is valid, a trivial effect of the manipulation with respect to that measure would provide strong evidence that the empirical defeat should be blamed on an auxiliary assumption rather than on the theory. The theory may be true or false, but the experiment fails to test it fairly. The failed manipulation check would provide the researcher with a good reason to try for either a different attitude manipulation, a different attitude manipulation check, or both. Just as an empirical victory need not render the theory correct, the empirical defeat would be an insufficient reason to discard the theory. Once the researcher has an attitude manipulation that passes a manipulation check, she can rerun the original experiment to test the theory. In that event, another empirical defeat would provide a much more convincing argument against the theory than sans the manipulation check.
We consider intervention research under the theory of reasoned action umbrella a useful case study for the importance of manipulation checks. Sniehotta et al. (2014) asserted that interventions based on the theory of reasoned action tend to fail to produce behavior change and concluded that this constitutes an important strike against the theory. However, there was little consideration of the possibility that perhaps the failure was not due to the theory being wrong, such as attitudes not influencing behaviors (through intentions), but due to the interventions failing to influence attitudes to begin with. If an intervention fails to influence attitudes, then the failure to produce eventual behavior change can hardly be fairly counted against the theory (St Quinton & Trafimow, 2022; St Quinton et al., 2021; Trafimow, 2015). Thus, again, manipulation checks are crucial for assessing empirical defeats.
Omission failures
Imagine a theory that intelligence increases task performance. An obvious preliminary test would be to measure intelligence and task performance and see if there is a nontrivial correlation coefficient between scores on the two measures. Suppose that the obtained correlation coefficient is near zero. One option would be to conclude that the theory is wrong. An alternative conclusion is that one or both measures are invalid. The researcher could perform additional studies to support or disconfirm the validity of the two measures. However, even if the researcher succeeds in convincingly supporting that both measures are valid, it may nevertheless be premature to conclude that the theory is false. And the reason concerns an auxiliary assumption that there are no crucial omissions.
To see the importance of this negatively stated auxiliary assumption, suppose that intelligence can influence task performance in two ways. One way is the theorized way, that intelligence increases task performance: IQ → TP+. The other way invokes the thus far omitted construct of boredom. It could be that intelligence increases boredom while performing the task, thereby decreasing task performance: IQ → B → TP-. Moreover, it is possible that both processes operate, so that one causal pathway, IQ → TP+, results in a positive effect of intelligence on task performance whereas the other causal pathway, IQ → B → TP-, results in a negative effect of intelligence on task performance. The net effect is near zero. It is unfortunate that correlation coefficients near zero are generally interpreted against the theory, an interpretation that has been famously formalized as a research prescription (Baron & Kenny, 1986). The possibility that constructs wrongly omitted from auxiliary assumptions are responsible for empirical failures is too rarely considered in psychology (Kline, 1998, 2015).
It is possible to elaborate. Of course, if a variable with causal status is omitted, then the researcher has no way to discover its role. But the foregoing point is that an endogenous variable can have more than one causal role, and the roles can conflict. If one or more of them involves a mediating but omitted variable, the researcher is unlikely to come to the right conclusion. Nor is this point restricted to purely correlational research. An experimental manipulation might also have multiple conflicting causal roles that cancel out to result in no ostensible result. Suppose the IQ example is modified so that participants receive training to make the task easier in the experimental condition, but not in the control condition. Participants in the experimental condition might be better at the task, which is a force for better task performance, and more bored, which is a force for worse task performance. As we saw in the correlational case, these contradictory forces could balance to result in a trivial net effect for the experiment.
A researcher who anticipates the possibility of a countervailing force through boredom might take measures to reduce the problem. One solution is to offer financial incentives for better performance, which potentially could instigate participants to devote full attention despite being bored, or might even reduce boredom. Then, too, the researcher could measure boredom and see if it correlates with IQ in the correlational example, or whether there is an effect of training on boredom in the experimental example. In general, there are many ways to handle the possibility of omitted variables. The problem is that many researchers fail to consider this possibility, thereby reducing the likelihood of addressing it.
Measurement validity
An obvious reason for empirical defeats is if one or more measures are invalid. Although validity of measurement is a huge topic, and there is no way to do justice to it here, it is nevertheless possible to address a small portion of that literature pertaining to the connection between a theoretical term and a corresponding empirical term by means of auxiliary assumptions.
Construct validity is probably most emphasized in the literature, and it refers to a matching of empirical and theoretical relations (Cronbach & Meehl, 1955). Construct validity can be characterized quickly by representing a theoretical relation with upper-case letters representing theoretical constructs, X causes Y or X → Y, and by representing an empirical relation with lower-case letters representing related empirical (observable) entities, x → y. The idea of construct validity is to show that the empirical relation, x → y, matches the theoretical relation, X → Y. In a true experiment, this can be done by using x as a manipulation of X, and using y as a measure of Y, and an empirical victory simultaneously supports (a) the theory and (b) that the manipulation and measure are valid. Or, in a correlational study, where x and y are both measures of X and Y, respectively, showing that x and y correlate supports X → Y and supports the validity of the measures, too. Thus, in the construct validity scheme, one supports the theory and validity of measures (or manipulations) simultaneously.
A limitation of construct validity is that it is possible for empirical relations to match theoretical relations for the wrong reason. For example, an attitude measure and intention measure might correlate not because attitudes cause intentions, but because mood does, and attitudes are correlated with moods. Thus, an incorrect auxiliary assumption can lead to a construct validity success.
More to the point of the present section, though, let us consider a construct validity failure. Although X is theorized to cause Y, there is no relation between x and y; the empirical relation mismatches the theoretical relation. There are at least three possibilities: the theory might be false, x might invalidly manipulate or measure X, or y might invalidly measure Y. At this point, before collecting any more data, it would be advantageous to ponder what the auxiliary assumptions are that link x to X and y to Y; why should we believe that x manipulates or measures X and y measures Y? Even without data, we believe that careful thinking would cast doubt on widely accepted measures. For example, measures of extraversion, one of the Big 5 traits, contain “enthusiasm” and “talkative” items. We could assume that both measure extraversion, but we could alternatively assume that the “enthusiasm” item measures enthusiasm and the “talkative” item measures talkativeness. Of course, extraversion aficionados can support the extraversion assumption with factor analyses, but that dodges a careful consideration of auxiliary assumptions. As we have seen, items can correlate for many reasons and loading on a factor should not be considered definitive, or even close to that. To ask a pointed question, “What are the auxiliary assumptions justifying that ‘enthusiasm’ and ‘talkative’ items measure extraversion?” We make no claim that there is no such justification, only that we have not heard one yet. Rather, researchers jump immediately to factor analysis and skip careful thinking about auxiliary assumptions. Thus, it is not surprising that psychology suffers from generally small effect sizes.
Staying with the extraversion example, suppose the theory is that being excited increases extraversion. A researcher has participants walk up a flight of stairs or not, and subsequently has them complete an extraversion measure but obtains a near null effect. Although it is possible that the theory is false, a careful look at the manipulation and measure suggests alternative possibilities. For example, although walking up a flight of stairs doubtless increases heart rate, it is not clear that this adequately stands in for “being excited.” It is possible that walking up a flight of stairs makes people tired, tense, irritable, and so on. Then, too, perhaps being excited increases talkativeness but not necessarily enthusiasm, and the extraversion measure confounds them, thereby creating a small effect where there should have been a large effect. Before concluding that the problem resides at the level of the theory, the researcher should address the other possibilities. This could include attempting different manipulations of “being excited,” different measures of extraversion, manipulation checks, and so on. Moreover, with respect to the potentially problematic extraversion, one simple check does not require further data. Suppose the researcher simply analyzes the “enthusiasm” item and “talkative” item separately and finds that the manipulation has little effect on the former but an impressive effect on the latter. Such a supplementary analysis would support a substitution of talkativeness for extraversion in the theory. Unfortunately, researchers rarely perform such analyses at the item level, which constitutes an opportunity cost. We believe that at least some validity questions would be effectively addressed, at no additional data collecting costs, if researchers would perform analyses at the item level rather than only at the level of the whole test.
Although we are not construct validity aficionados, and agree with Slaney (2017) that there is much confusion among researchers about what construct validity means, we believe that some construct validity criticisms are overstated. Maraun (2012) argued that valid measurement is about following particular rules and has nothing to do with results. Consider the following quotation:
That men are, on average, say, 167 pounds, has no implication for what it is to correctly measure mass nor, certainly, whether the numbers on which the claim is based are, in fact, measurements of mass. On the contrary, the claim is not a claim about mass at all unless the numbers are masses, hence, unless they were taken in conformity with the rules for the measurement of mass. (p. 82)
The quotation reveals important lessons about why it is insufficient to merely recognize the cruciality of auxiliary assumptions and that one should recognize, too, the myriad implications. For example, the quotation confuses mass and weight. The man who weighs 167 pounds at sea level would not weight 167 pounds on the top of a mountain, on the moon, and so on. But that man would have the same mass. Thus, a measure of weight is not a measure of mass unless auxiliary assumptions are made linking unobservable mass to observable weight. The strong implication is that measurement, even in physics, is highly dependent on auxiliary assumptions or on one’s measurement model that contains these auxiliary assumptions.
Secondly, even in physics, validity is not only about rules of measurement, with findings irrelevant. Suppose a physicist weighed the man, at sea level on earth, and found the weight equal to zero. The physicist could question the theoretical value of mass, but the physicist could alternatively question whether the scale is working. Taking the latter route might cause the physicist to get another scale. If the latter scale gives a reasonable reading, that would support that the measuring device (the scale) did not work correctly. Thus, we see that empirical findings do have a role to play in assessing the validity of measurement; validity is not only about following rules.
Returning to psychology, suppose a measure of extraversion did not correlate with desire to engage in conversations, desire to interact with people, and so on. Although the present authors question whether there is an entity to which extraversion refers (we suspect Maraun might agree with us here), if we were to believe that extraversion is real, the contrary findings would render it perfectly reasonable to suspect that one or more auxiliary assumptions connecting extraversion to the measurement items is faulty.
Conclusion
The issue of how much to attribute credit or blame to theories versus auxiliary assumptions, for empirical victories or empirical defeats, has long been a lively philosophy of science issue (Duhem, 1914/1954; Whewell, 1840). It is unlikely that any single article could definitively solve the issue either for science, generally, or for psychology, specifically. Nor do we claim to have accomplished that here. Our aim was more modest. In the event of empirical victories, we discussed some traditional and nontraditional techniques for aiding the process of attributing credit, and in the event of empirical defeats we discussed some techniques for aiding the process of attributing blame. Either way, the techniques were not definitive and were accompanied by caveats or qualifications. When it comes to theory-testing, there are rarely easy answers.
That each technique is hedged with caveats or qualifications is not a reason for pessimism. It is difficult to soundly attribute 100% credit or blame to theories or auxiliary assumptions, but that difficulty should not be allowed to obscure that sometimes researchers are able to make a convincing case. The range between complete belief or disbelief in a theory takes in much territory, and if a research program can push the balance of evidence substantially in favor of or against a theory, that constitutes an important contribution. For instance, it is possible that despite the massive evidence in favor of evolutionary processes, creation nevertheless could have occurred. Evolutionary processes cannot be proven beyond any doubt whatsoever, but that tiny doubt should not be allowed to degrade the contributions made by researchers who have gained relevant evidence rendering evolution much more plausible than creation (see Johnston, 1999 for a review). That evidence is extremely convincing, even if it does not constitute absolute proof. It is possible, but not plausible, that there is another planet like Earth somewhere else in the galaxy, and a being with advanced powers created one for themselves just like it, and we call that created planet “Earth.” Thus, the massive evidence for evolutionary processes on the original planet would have been reproduced with respect to Earth, despite Earth having been created. That this improbable scenario remains possible demonstrates the futility of insisting on absolute proof, but science can progress sans absolute proof.
There exist clear psychology examples where auxiliary assumptions transformed empirical defeats into empirical victories, thereby shifting the balance of evidence from substantial disfavor to substantial favor. Our favorite stems from the famous review of the attitude literature by Wicker (1969). Attitudes had long enjoyed dominance as the most important construct in social psychology (Allport, 1935), but Wicker’s review showed that attitudes were only slightly correlated with behaviors. The myriad empirical defeats Wicker reviewed contradicted the importance of attitudes in causing behaviors.
Fishbein and Ajzen (1975) addressed the empirical defeats at the level of both theory and auxiliary assumptions. At the level of theory, they conceptualized attitudes as towards behaviors rather than towards objects. This change facilitated the assertion that attitudes determine behavioral intentions which, in turn, determine behaviors. In addition, Fishbein and Ajzen argued that there is a normative, as well as attitudinal, route to behavior. More important, however, for rendering salient the importance of auxiliary assumptions, they argued that the reason attitudes performed so badly is that previous researchers measured them invalidly. To measure attitudes validly, Fishbein and Ajzen insisted that behaviors have four elements: action, target, time, and context. Attitude measures must correspond with behavior measures with respect to all four elements. When researchers commenced measuring attitudes using the new auxiliary assumptions pertaining to attitude measurement, impressive correlation coefficients became the rule rather than the exception, as the meta-analysis by Kraus (1995) eventually demonstrated. Then, too, the impressive correlation coefficients were taken as empirical victories for the theory, thereby reversing the empirical defeats for the importance of attitudes that Wicker (1969) had reported.
In conclusion, there is reason for optimism despite the difficulty in proving or disproving theories in absolute terms. There are many possible gradations of the weight of the evidence, and a careful consideration of auxiliary assumptions can crucially inform researchers’ assessments. Assessments of the weight of the evidence should be nuanced, with the caveats and qualifications we reviewed playing a nontrivial role.
Footnotes
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
