Abstract
The central roles of science in the field of remedial and special education are to (a) identify basic laws of nature and (b) apply those laws in the design of practices that achieve socially valued outcomes. The scientific process is designed to allow demonstration of specific (typically positive) outcomes, and to assist in the attribution of those outcomes to controlled variables. Although growing recognition is being given to the importance of replication in this process, equal consideration should be given to the function of publishing studies that document negative (or null) results. In this manuscript, we outline the features of negative results in educational and psychological single-case intervention research. We also discuss the assessment, methodological, and statistical dimensions of negative results that should be considered when reporting negative results. The importance of replication studies (direct, systematic, and clinical) is also discussed within the context of negative-results research.
Introduction
Empirical knowledge regarding the effectiveness of interventions in the social and educational sciences is generated through publishing and summarizing research results. Traditionally, studies have been published and subsequently summarized that report positive findings, or positive results (e.g., results that document an experimental effect). Less emphasis has been placed on publishing negative findings, or negative (or null) results (e.g., results that do not document an experimental effect). In randomized controlled trials (RCTs), negative results typically refer to statistically nonsignificant differences between or among groups that receive different treatments, interventions, or conditions. Similarly, negative results can occur in single-case design (SCD) experiments when there are no documented differences (visually and/or statistically) between baseline and intervention conditions. Negative results are not to be confused with negative effects, or undesirable side effects of the intervention that occur as a function of the intervention. Essentially, the term negative results should not carry a connotation of unimportance as finding no effect of the intervention can be of great importance in applied and clinical research. Our focus in this introductory article, and in this Special Issue, is on negative results in SCDs.
We dedicate this article and this Special Issue to William R. Shadish (1949–2016). A wonderful colleague and friend, Will was a champion of experimental research in psychology and education and of the criticality of publishing negative results in those fields. In his own characteristically engaging fashion, he often argued that the exclusion of negative-results studies creates the very real potential for publication bias, a false impression of a scientific “finding,” and practice misdirection.
Will made significant contributions to the special panel formed by the What Works Clearinghouse (WWC) to create pilot standards for SCD research. The resulting White Paper (Kratochwill et al., 2010) and a subsequent paper published in this journal (Kratochwill et al., 2013) distinguished between design standards and evidence criteria, a point that features the importance of negative results and which is the focus of this article. The WWC panel was aware of the documented problems associated with academic journals publishing primarily positive results in intervention research (e.g., Dishion, McCord, & Poulin, 1999; Greenwald, 1975; Ioannidis, 2005; Rosenthal, 1979) and of the importance that negative results have for establishing the scientific foundation for interventions in the social and educational sciences (e.g., Ferguson & Heene, 2012; Kazdin, 1998; Kratochwill, Stoiber, & Gutkin, 2000; Kupfersmid, 1988; Lykken, 1968).
With Will’s encouragement to include negative results in the identification of scientific knowledge, and the active support of the Remedial and Special Education (RASE) senior editors, this Special Issue emerged. Our goals here are to (a) define the value of negative results for basic and applied research in education and psychology and (b) provide examples of empirical research that makes a fundamental contribution to the field through reporting negative results. Will’s contributions to the betterment of educational research methodology and analysis are immeasurable. Regarding negative results, in particular, portions of this article stem directly from Will’s input to heretofore unpublished collaborative work that we had initiated. We are certain that Will would be pleased to know that with this article, his concerns about the negative consequences of negative results will live on―and might even have a positive influence on the quality of the SCD intervention literature in the future.
The Importance of Negative Results in Evidence-Based Practice
The current emphasis on evidence-based practices requires defining where, when, and with whom practices are and are not effective (Flay et al., 2005; Horner & Kratochwill, 2011). A comprehensive science that meets applied needs must include an analysis of (a) specific conditions for which practices are effective, (b) the conditions for which practices are not effective, and (c) any clinically undesirable effects of the practices. The issue of negative results is especially important in the evidence-based practice movement, where a body of literature is used to support or not support an intervention. As we have noted, negative results can occur in SCD research as well, although these have less often been discussed within the context of research on evidence-based interventions. In fact, in an entire issue of Perspectives on Psychological Science devoted to replicability in psychological science, and in which the topic of negative results was discussed at great length (Pashler & Wagenmakers, 2012), not one contributor discussed the role of SCD in the so-called “crisis of confidence.”
Examples of negative-results research in SCD experiments is less well described than the literature on RCTs but, for several reasons, remains highly important in the development of evidence-based practices in psychology and education (Kratochwill et al., 2000). First, in practice, we want to promote interventions that are effective for the problem under consideration. Some interventions are selected for individuals that have never been tested, or when they are tested they demonstrate no improvements on selected outcome measures. Consider the example of sensory interventions that have been demonstrated to be common practice for children with developmental disabilities. Yet, when interventions derived from sensory integration theory are subjected to rigorous experimental tests in both group and SCD research with developmentally disabled children, evidence documenting positive outcomes has not materialized (e.g., Barton, Reichow, Schnitz, Smith, & Sherlock, 2015).
Second, it would be most desirable for researchers to be open to providing a fair and unbiased test of the intervention in an experimental trial, a reason that led us to support the distinction between design standards and evidence criteria in SCD research (Kratochwill et al., 2013). Thus, if the researcher embraces high-quality design standards in SCD intervention research, negative results should be an acceptable evidence criterion because a rigorous test of the intervention would have been conducted. In fact, we would argue that in some cases negative results may be a larger contribution than positive results for scientific understanding of an intervention. That is often the case when attempting to understand nuances about an intervention or the scope of its influences, including, for example, examining an intervention’s generalizability across participant populations and contextual variations (e.g., Bracht & Glass, 1968). One vision of how a “negative results” process-and-policy rationale might unfold with scientific journals is provided by Kittelman, Gion, Horner, Levin, and Kratochwill (this issue).
Third, a seemingly attractive intervention might be extremely expensive to produce and/or costly in terms of implementation time and logistical factors. Then, when a rigorous experimental test of the intervention is conducted, it may have no substantial impact on participant outcomes (i.e., negative results are produced). Such a finding is of “practical significance” because it may lead researchers to redirect their financial and human resources to alternative interventions that are more likely to have a positive impact on the specified outcomes.
Some Policy Issues Surrounding Negative Results
It is no secret that the majority of intervention-research publications in the scientific literature consist of studies reporting positive results. After all, researchers are interested in successful outcomes of interventions so that, ideally, they can be adopted into practice to improve valued outcomes, including the quality of life of individuals. In fact, our professional journals may have policies that promote positive results when publishing empirical work on the effectiveness of interventions (see Kittelman et al., this issue). However, the infrequent publication of negative-results studies (relative to publication of positive-results studies) leads to an incomplete and potentially misleading literature corpus. This situation, in turn, produces the file drawer problem where important, but negative, results” are never submitted for publication (Rosenthal, 1979)―an issue that has been documented to be a concern for publication policy practices (e.g., Shadish, Doherty, & Montgomery, 1989). It could also be argued that ignoring negative results is part of the culture in publication policies (Nosek, Spies, & Motyl, 2012) and may even be on the rise (Fanelli, 2012). Such a scientific culture has created a “publication bias” (e.g., Ferguson & Heene, 2012; Greenwald, 1975), and it is increasingly recognized that this circumstance has deleterious effects on science (Pashler & Wagenmakers, 2012), could promote academic misconduct (Fanelli, 2012), and, notably, limit the development of a knowledge base for evidence-based practices and policies (Ioannidis, 2005; Kratochwill et al., 2000). The file drawer problem and editorial bias operate in tandem to generate the misleading literature predicament in that (a) researchers are less likely to submit for publication studies that yield negative results and so such studies end up in researchers’ file drawers (or nowadays, in the dark recesses of their computers’ hard drives) and (b) the (smaller number of) negative-results studies that are submitted are less likely to receive recommendations to be published from reviewers and editors (see also Kittelman et al., this issue).
Publication bias and negative results in SCD research
Publication bias refers to the practice of publishing studies that possess certain characteristics that make them more likely to be accepted for publication (Ferguson & Brannick, 2012; Song, Easterwood, Gilbody, Duley, & Sutton, 2000). If a group of studies examining the effects of a practice includes some studies documenting an effect and some studies that do not document an effect, but the former are more likely than the latter to have been published, the field will receive a biased view of the practice. There is meta-analytic evidence that publication bias exists in education and psychology journals (e.g., Polanin, Tanner-Smith, & Hennessy, 2016) and specifically in special education journals (Gage, Cook, & Reichow, 2017). Generally, the features that lead to publication bias include reviewers’ preference for studies with statistically significant effects and large (i.e., practically significant) effects, as well as studies with rigorous experimental methods (Sutton, Duval, Tweedie, Abrams, & Jones, 2000). (We note here that the present authors do not regard publication that favors the latter class of studies not as a “bias” but as a critical defining characteristic of publishable studies [see also Kittelman et al., this issue]). In the domain of SCD research, publication bias exists, but we know less about its form or functioning (Kilgus, Riley-Tillman, & Kratochwill, 2016). For example, some parts of the SCD research community may have an explicit ethos that an intervention must produce a functional relationship (essentially equivalent to an unambiguous or strong intervention effect) for a study to proceed through the publication process. If so, publication bias in SCD research is typically driven by finding/producing effects rather than by the common “statistical significance” criterion that is characteristic of conventional “group” research. Yet, generally unexplored is the role that publication bias plays across a number of different SCD domains (e.g., in the current research literature, perceptions of journal editors, researchers, and reviewers), as well as the possible role that publication policies regarding negative results can play in developing evidence-based practices. Interestingly, publication bias maybe driven, in part, by researchers who make the decision not to submit a negative-results paper for publication (Cook & Therrien, 2017; Franco, Malhotra, & Simonovits, 2014; Olson et al., 2002) despite recently reported evidence that journals may be increasingly willing to publish such results (e.g., Driessen, Hollon, Bockting, Cuijpers, & Turner, 2015).
Traditionally, scientific journals have been unlikely to publish negative results (e.g., Atkinson, Furlong, & Wampold, 1982; Mahoney, 1977), including replication research (Neuliep & Crandall, 1993)―see Cook and Therrien (2017). Some researchers have explored the relationship between negative results and publication bias in SCD research (Shadish, Zelinsky, Vevea, & Kratochwill, 2016; Sham & Smith, 2014). Sham and Smith (2014) assessed publication bias by comparing effect sizes in SCD research in published studies (n = 21) and nonpublished dissertation studies (n = 10) in the area of pivotal response treatment with a nonoverlap method called percent of nonoverlapping data (PND). The authors reported that PND for published studies was 22% higher than for unpublished studies. In the Shadish et al. (2016) study, SCD researchers were surveyed about their publication practices and results suggested that researchers expressed a preference for submitting manuscripts for review that showed large effects. Similarly, journal reviewers expressed a preference for large effects when making their publication recommendations. Although data on SCD publication practices is limited, the data of which we are aware suggest that there is likely a preference for positive-results studies and very likely a publication bias in the SCD intervention literature.
Negative results increasingly at the forefront of policy changes
Although targeting when and where research innovations do not work is an underdeveloped topic in SCD research, the topic is increasingly being examined in the scientific literature and has begun to affect research policies. First, and as was noted above, the WWC Single-Case Design Pilot Standards were developed specifically with the distinction between design standards and evidence criteria, thereby suggesting that if a fair test of an intervention results in negative results, then such results should be welcomed into the scientific literature (Kratochwill et al., 2010, 2013). And, as we have mentioned earlier, an increasing number of literature reviews of intervention research have applied this distinction and have included SCD studies in which negative results were found (e.g., “Functional Behavior Assessment,” 2016; Kiuhara, Kratochwill, & Pullen, 2017; “Pivotal Response Training,” 2016).
Second, the U.S. Department of Education advanced the importance and interpretation of negative results in their policy communications. For example, the Institute of Education Sciences (IES) noted the importance of negative results in its 2013 Webinar on Administrative Requirements, where it committed to “identifying what does not work and thereby encourage innovation and further research.” More recently, the Department produced a report in which the “no effect” outcome was presented along with possible reasons for such findings (Seftor, 2016). Possible reasons advanced for negative results included (a) the outcomes may represent a problem with theory in the translation to development and testing of an intervention, (b) a failure of implementation, and (c) the design could not assess the intervention with precision. Although Seftor (2016) did not feature SCD research in his discussion of negative results, each of the reasons that he outlined for conventional group research can be generalized to negative-results findings in SCD research. In concert with our perspectives, we are suggesting that “negative but informative” results can be distinguished from “negative but flawed” results. We are not arguing that all negative results should be reported (e.g., when a study does not meet high-quality research design standards, does not document sound intervention fidelity, or focuses on an intervention that has wide-scale adoption but no empirical support). Rather, we are proposing that SCD studies investigating interventions that are well designed, well implemented, and societally valued should be published, regardless of the experimental outcomes (see, for example, Kittelman et al., this issue; Rosenthal, 1966; Walster & Cleary, 1970).
Third, there is growing recognition of the importance of negative results in the intervention-research literature, as witnessed in this and other Special Issues in special education journals (Cook & Therrien, 2017). This recognition stems from both a commitment to high-quality scientific knowledge, and ethical concerns in which withholding publication on negative results violates a basic tenet of the scientific process. Writing in Behavioral Disorders, Cook and Therrien (2017) reviewed websites for 39 major journals focused on special education, in which they found that only one gave guidance to authors about the publication of negative results.
Dimensions of Negative Effects, Selective Results, and Erroneous Results
Beyond domain of negative results, there are features of intervention research that may overlap with negative results and may largely be responsible for observing no effect of the intervention. These features include (a) presenting research in which outcomes consist of negative effects, or undesirable side effects of the intervention that are documented to occur―and again, which should not be confused with our present focus on negative results. The features also include (b) selective results, or the reporting of desired positive findings while withholding contradictory results and/or information about unwanted side effects; and (c) problems that emerge from erroneous results, or the reporting of results that ostensibly support an intervention’s effectiveness but do not hold up to scientific scrutiny as a result of critical methodological and/or data-analysis shortcomings (e.g., Levin, 1985). Because, to date, little has been studied and publicized about the impact of negative effects, selective results, and erroneous findings on the scientific knowledge base for evidence-based practices in the area of SCD intervention research, there is a need for researchers, journal editors, and reviewers to be better informed about the influence of such features when making decisions about the review and publication of studies based on this research genre.
Negative Results Versus Negative Effects in SCD Research
As was just noted, an important distinction must be made between the present focal topic of negative results/findings and negative effects (or what are sometimes referred to as iatrogenic effects). Negative effects of interventions producing adverse side effects on participants should be documented―an issue that Barlow (2010) discussed in the psychotherapy literature. Negative effects of an intervention can occur with interventions that produce positive or negative results and should be presented in scientific reports. Nevertheless, like negative results, negative effects on individuals exposed to interventions are unlikely to be published (Daves, 1994; Glass & Smith, 1978), thereby leading further to publication bias in our journals. For both the good of science and the good of practice, it is paramount that SCD researchers present all effects produced by their interventions in published studies irrespective of whether the results are positive or negative.
There are several possible dimensions of negative effects in SCD intervention research. First, participants in the experiment may actually deteriorate or get worse after receiving the intervention. Such a finding may include not just a negative data-analysis result but also a negative (potentially harmful or aversive) effect on the study’s participants. As an illustration, such a circumstance may occur in a SCD multiple-baseline design study across participants in which one or more of the participants exhibits deterioration in a particular condition after receiving the intervention. This “mixed result” finding should lead the investigator not only to present a negative-effect outcome and attempt to document the reasons for it but also to include plausible strategies or alternative interventions to ameliorate the affected participants’ condition.
Second, negative effects may also occur when an investigator finds that the intervention not only produced positive or even null results on the targeted outcome measure but also caused unanticipated problems to emerge among participants. For example, an investigator may target and find improvement in academic outcomes following the intervention, only to discover that some participants also demonstrated increased behavior problems. Such a circumstance should lead the investigator to address the issue in the subsequent scientific report, which might also be expanded to include a modified follow-up experiment. As a further illustration, negative effects may be documented in SCD studies in which the researcher is administering a potent drug intervention (Barlow, Nock, & Hersen, 2009). The negative effects may be due to the drug composition, dosage level, duration, interaction with another medications, and so on.
Finally, even if an intervention proves to be of benefit to individuals (i.e., it yields positive results), an unwanted negative effect of the intervention would occur if social validity criteria demonstrate that the participants report that they did not like the intervention and would not use or implement it in the future. In such circumstances, social validation measures might include subjective evaluation of the intervention, suggesting that acceptability is low under certain conditions such as time to implement, the presence of objectionable content, logistical factors, and the like (see Kazdin, 2011).
Selective Results
Selective results refer to the willful withholding of any findings in a single study or in subsequent replication attempts (i.e., a series of SCD investigations in which the intervention is repeated in the same or in independent studies)―(see also our discussion below for selective-results issues in replication attempts). Consider the following possibilities for selective results reporting. An investigator might conduct several studies in which the first few experiments in the series do not demonstrate positive effects (i.e., the studies yield negative results) or only modest effects that are attributable to the intervention. Subsequent studies in which the investigator “tweaks” the intervention may produce positive findings. Given the culture of publication bias, the scientific community may learn only of the positive-outcome studies insofar as they are the ones that are more likely to be accepted for publication and subsequently reported in the literature. Selective results may also appear in cases where multiple-outcome measures are included in a single investigation. Some of the measures may demonstrate positive outcomes, whereas others demonstrate no positive outcomes or even outcomes in the direction opposite from what was predicted. However, the investigators may report only the outcomes that appeared in a hoped-for (or eye-catching) way. Exacerbating the selective-reporting problem, some SCD researchers have indicated that they may drop cases with small effects prior to submitting the manuscript for publication consideration (Shadish et al., 2016).
An example of this type of selective reporting was spotted by Engber (2013) in relation to two concurrently published articles that were based on the same 47,000-participant data set (from The National Runners’ and Walkers’ Health Study) concerning the effects of walking versus running on participants’ health risks and benefits. In one article (Williams, 2013), where the researcher’s purported intent was to “test whether equivalent changes in moderate (walking) and vigorous exercise (running) produce equivalent weight loss under free-living, non-experimental conditions,” it was found that changes in the participants’ body mass index were “significantly greater for running than walking” (p. 706). In the second article (Williams & Thompson, 2013), where the purported intent was to “test whether equivalent energy expenditure by moderate-intensity (e.g., walking) and vigorous-intensity exercise (e.g., running) provides equivalent health benefits,” the authors concluded that “[e]quivalent energy expenditures by moderate (walking) and vigorous (running) exercise produced similar risk reductions for hypertension, hypercholesterolemia, diabetes mellitus, and possibly [coronary heart disease]” (p. 1085). As Engber put it, “So there you have it, and there you don’t. Running is better for your health, or perhaps it isn’t.” However, the issue of concern here is that one article mentions nothing about the findings reported in the other article, which can be regarded as questionable (selective-reporting) behavior on the part of the researchers. Simmons, Nelson, and Simonsohn (2011) addressed this selective-reporting issue by proposing that authors must list all variables conducted in a published study (Requirement 3, pp. 1362–1363)—see also Maxwell and Kelley’s (2011, pp. 172–176) discussion of various selective data-analysis and results-reporting practices. Clearly, with respect to the reporting of selective results, major ethical implications need to be considered.
Erroneous Results
Another potential investigator issue of concern is that of incorrect data management, analysis, and interpretation, or erroneous results. In SCD intervention research, there are typically two options for analyzing and reporting results of the experiment, visual analysis, and visual analysis supplemented with statistical analysis. The outcomes of the experiment comprise the focus of the WWC Pilot Standards’ “evidence criteria” and are considered separately from the “design standards” (Kratochwill et al., 2013). Historically, evidence criteria in SCDs have been typically assessed through visual analysis of the data. However, erroneous results can emerge when there is unreliability in the visual-analysis process, a feature that has been documented over the years―See Kratochwill, Levin, Horner, and Swoboda (2014) for a review of visual-analysis approaches and concerns. Given our previous discussion about reviewers and editors favoring the publication of studies that report positive results, it should come as no surprise that through visual-analysis investigators tend to “see” in their data effects that are in reality negative results. If investigators are claiming that an observed effect is “real” when in fact it is not (i.e., it is a negative result or a chance finding), then, in inferential statistical-analysis jargon, they are committing a Type I error―and an abundance of Type I errors is exactly what has been found in studies of the visual-analysis process (Kratochwill et al., 2014). So, just as in the statistical-analysis Type I error literature, negative results are likely masquerading as positive results through the inaccurate or unreliable visual analyses that are typically conducted in SCD research. Fortunately, guidelines for conducting visual analysis are available and various visual-analysis training protocols have been developed to improve the visual-analysis process (see Kratochwill et al., 2014).
The Importance of Publishing Single-Case Negative-Results Research
There are several compelling reasons to consider negative results in SCD intervention research. First, we need to document that specific interventions that are touted as effective are actually not effective when given a stringent test (e.g., Barton et al., 2015). Second, we need to identify interventions that are effective in some (but not all) contexts, thereby improving the precision with which we can state where, when, and how an intervention is likely to be effective (Kilgus et al., 2016). Third, we need to identify within-subjects intervention conditions and circumstances that do and do not produce positive outcomes, such as when an ABABCBC design demonstrates a positive effect when Condition B is introduced but not when Condition C is added (Barlow et al., 2009). Nevertheless, in SCD research, the failure to find differences among experimental conditions may be attributable to factors that are not easily discernible. These factors need to be addressed explicitly by researchers, reviewers, and journal editors when a negative-results study is considered for publication (see Kratochwill et al., 2000; Levin, 1985; Kazdin, 1998; Tincani & Travers, IN PRESS, for a review of these factors).
In this section, we discuss certain aspects of research reporting that can help in the publishing of negative-results research. These aspects include (a) detecting different intervention effects on different outcome measures, (b) clarifying and elaborating the contextual variables under which negative results are detected, and (c) examining the dimensions of replication research.
Exploring Variable Effects on Different Outcome Measures
Variations in outcomes across multiple dependent variables can assist researchers in understanding the influence of the intervention under consideration and may even lead to theoretical insights (Kazdin, 1998). Similarly, and as an extension of Campbell and Fiske’s (1959) discriminant validity notions, in SCD studies that incorporate both multiple interventions and multiple-outcome measures, predicting and detecting differential intervention effects (essentially intervention-by-outcome interactions) can be extremely valuable from both theoretical and practical perspectives. As Levin (1989) indicated,
The basic philosophy underlying this research approach is simply that when a particular independent variable differentially affects two or more dependent variables in ways that can be specified on an a priori basis, much more is learned about the independent variable’s operation than is learned either when only a single dependent variable is affected or when two or more dependent variables are affected in the same manner. (p. 86)
An important implication of this philosophy is that SCD researchers should not only include a broad range of outcome measures in their investigations to test adequately the scope or impact of their interventions (to assess the interventions’ external validity), but they also should include outcome measures that are differentially sensitive to different interventions or intervention variations (to assess the interventions’ discriminant validity). Obviously, decision rules are needed regarding sensitivity of the measures, but a variety of indicators might be considered (e.g., statistical, clinical, social validity―see Ogles, Lambert, & Masters, 1996).
Specifying Conditions Under Which Negative Results Are Detected
In SCD intervention research, the study can often be designed to demonstrate specific conditions under which positive and negative results are likely to occur (Kazdin, 2011). Technically, the finding of an equivalence of a novel intervention and an existing evidence-based intervention can be conceptualized as a no-differences finding. For example, consider the three classes of comparative-intervention SCDs where such a finding is possible. First, intervention comparisons can be scheduled among different conditions as in the Alternating Treatment Design, wherein the investigator compares two or more interventions with each other (see Kratochwill et al., 2013). It is possible that some comparisons would show equivalence, as has been found in studies by Reichow, Barton, Sewell, Good, and Wolery (2010) who systematically examined the effects of wearing weighted vests on children with developmental disabilities. Reichow et al.’s study, along with a series of additional investigations, all of which yielded negative results, provide a serious challenge to weighted-vest advocates whose empirical research support is derived from a single precedent study with questionable positive outcomes (Fertel-Daly, Bedell, & Hinojosa, 2001; and see Barton et al., 2015, for a comprehensive review).
Another example of where negative results might be demonstrated is in the use of complex phase-change within-series designs (Ingram, Lewis-Palmer, & Sugai, 2005). In such designs, for example, ABABA(B + C)A(B + C), different intervention components are systematically introduced and assessed. An intervention package with one element (B) may be unsuccessful; however, when two (B + C) or three (B + C + D) components are incorporated into the package, the intervention effects may be strong. Third, in a Multiple-Baseline Design across participants, settings, or behaviors, it is possible that in one of the multiple series the anticipated change does not occur. Thus, the researcher may still meet WWC design standards even though the intervention was found not to be effective in one of the series.
Demonstrations in SCD research of no-difference findings can provide decisions about how to increase the strength of an intervention in replication-research attempts (see discussion below). They may also provide a refined understanding of why a particular intervention did not work effectively in a particular context. Thus, the implications of obtaining negative results within this SCD framework need to be considered. In addition, discussion of no-difference findings within the context of either “additive” or “dismantling” intervention strategies needs to be made explicit with respect to evidence-based interventions (Barlow et al., 2009).
Conducting and Reporting Replication-Research Attempts
Virtually all areas of the social sciences and education emphasize the replication of research findings. Although with foci somewhat different from those in our discussion of replication attempts in SCD intervention research, replication is critical in special education research (Travers, Cook, Therrien, & Coyne, 2016). The special issue of Perspectives on Psychological Science explicitly included replication as a central feature of establishing the beneficial nature of psychological science (Pashler & Wagenmakers, 2012; see also Winerman, 2013). Concern for replication has been regarded as so important in psychology that the Reproducibility Project was established with the distinct goal of assessing the rate and predictors of reproducibility in scientific research (see the Open Science Framework’s website at http://openscienceframework.org/). More recently, the editor of the American Psychological Society’s flagship scientifically grounded-research journal, Psychological Science, has been soliciting and publishing what are called preregistered direct replications (PDRs), which
should reproduce the original [study’s] methods and procedures as closely as possible, with the goal of measuring the same effect as in the original study . . . Some direct replications will report ‘successes’ in which the original findings are closely replicated, and some will report unambiguous failures to replicate (made compelling by fidelity to the original studies, high statistical power, and appropriate analyses). Both of those outcomes are valuable and informative. (Lindsay, 2017, pp. 1191, 1192)
Within the context of SCD replication research, negative results play a prominent role. Reliability of the findings across independent investigators and investigations is critically important for determining intervention efficacy, even though replication itself is associated with a number of possible ambiguities that can occur―such as selective methodological decisions, investigator expectation and bias, and sequential sampling, among others (for elaboration, see Francis, 2012, and the rest of the 2012 special issue of Perspectives on Psychological Science).
There are also examples of interventions that have not been documented to be effective and yet receive positive press and high usage. For such instances, documentation of negative effects may be very valuable. As is nicely illustrated in the debate about the efficacy of facilitated communication, the American Psychological Association (APA) issued a policy statement about its ineffectiveness despite vocal advocates (http://www.apa.org/research/action/facilitated.aspx). Finding that results do not reproduce over time may lead to revisions in the intervention, restrictions on its generalizability, or possibly the need for strengthening the intervention―such as linking it to other interventions of known beneficial qualities (Shadish, Cook, & Campbell, 2002). The WWC Pilot SCD Standards state emphatically that the validity and credibility of intervention findings are enhanced when reproduced across different cases, studies, and research groups (see also Horner et al., 2005). In this regard, the Standards’ 5-3-20 “rule” was advanced for including SCD experiments in a systematic evidence review, with the ultimate goal of designating an intervention as an evidence-based practice (see Kratochwill et al., 2013).
The 5-3-20 rule has increasingly been applied in numerous literature reviews (see Kiuhara et al., 2017). For domain-specific examples, see Maggin, Chafouleas, Goddard, and Johnson (2011) for a review the literature on token economies as an intervention for students demonstrating behavior challenges and Maggin, Briesch, and Chafouleas (2013) for a review of the literature on self-management. Nevertheless, the 5-3-20 rule is in need of further development and refinement, especially with respect to the type of replication research that is to be conducted and eventually included in literature reviews that inform evidenced-based practices. Moreover, the 5-3-20 rule does not systematically address negative results or speculations about why an intervention was or was not effective. In the future it may be useful to add a fourth criterion for designating evidence-based practices, namely, as related to the paucity of published negative results.
Within the context of SCD replication-research attempts, three types are considered critical: direct replication, systematic replication, and clinical replication (Hersen & Barlow, 1976). Conceptual and methodological issues surrounding negative results are different across the different types of replication studies and so the findings would be interpreted somewhat differently depending on the type of replication study conducted. Space does not permit discussion of these three replication types here but see Barlow and Hersen (1984), and Travers et al. (2016), for a more extensive consideration of them.
Summary and Conclusion
The role of negative results has at least three important contributions to the emergence of scientific knowledge about evidence-based practices in education and psychology. First, negative results document proposed hypotheses and interventions that do not contribute to improved outcomes (i.e., application of the proposed practice does not produce anticipated improvement in valued outcomes). Documentation of negative results may be of greatest value in conditions where an educational or clinical procedure is highly touted, but insufficiently examined. Often, research demonstrating negative results also includes assessment of contrasting strategies or procedures that do produce desired change in valued outcomes.
Second, a more nuanced contribution of negative results is to refine the conditions under which an innovation or intervention is and is not effective (e.g., across populations, contexts, implementers). No intervention produces positive outcomes for all participants in all contexts. Defining the conditions under which improvement in valued outcomes is and is not likely will be an ongoing contribution of applied research, and in replication research, in the fields of education and psychology.
Third, negative results may document situations where an innovation or intervention is “iatrogenic” (i.e., is functionally related to effects in the opposite direction of what prior research would have predicted or where it demonstrates negative effects, as was noted in this article). There are cases where a particular innovation or application has strong conceptual or social value, but unintentionally results in detrimental or negative side effects. Rigorous demonstration of iatrogenic effects is helpful to a larger understanding of basic laws of nature, and approaches for developing clinical technology.
We contrast the above roles of negative results from the simple absence or “disconfirmation” of anticipated positive effects due to procedural insufficiencies. Rigorous, methodologically sound, intervention research is challenging to conduct. It is certainly possible for a study to be conducted without reliable measurement of the dependent variable. However, if the valued outcome(s) of a study cannot be monitored with accuracy and precision, then any findings become difficult, if not impossible, to interpret. Similarly, any educational or clinical study carries a substantial burden to document that the intervention or innovation under analysis (i.e., the independent variable) was manipulated as intended (i.e., with fidelity), in the context proposed, and by and with the individuals stipulated. The growing expectation that such studies measure up to the “integrity of intervention” is indicative of this concern.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
