Abstract
Positive behavior interventions and supports (PBIS) is a three-tiered framework shown to improve student behavioral outcomes. There has been considerable investment in scaling up Tier 1 PBIS but less focus on scaling Tier 2 supports. Having greater insight into teachers,’ students,’ and families’ perceptions regarding the social validity of Tier 2 interventions may further facilitate the scale-up of this more resource-intensive level of the multitiered model. The primary objective of this systematic literature review was to evaluate the social validity of Tier 2 interventions in 48 experimental studies, 65% of which systematically assessed social validity. Almost all studies reporting social validity data surveyed student participants and teachers (94%). Results indicated teachers and students primarily reported positive perceptions regarding Tier 2 interventions, although some teachers expressed concerns with effectiveness of the outcomes. Findings provide insights that can help promote the scale-up and broader dissemination of Tier 2 supports.
Challenging behavior adversely affects students’ academic achievement, peer relationships, and self-esteem and can lead to school suspension and school dropout (Noltemeyer et al., 2015; Preeti Kumar et al., 2016; Tetzner et al., 2017). To prevent and address challenging behavior before it leads to negative outcomes, many schools are implementing positive behavior interventions and supports (PBIS). PBIS is a three-tiered framework that employs evidence-based practices with increasing intensity based on student needs as determined by ongoing assessment data with the goal of supporting all students’ behavioral needs (Maggin et al., 2016; McCurdy et al., 2016).
A growing body of research indicates PBIS has positive effects on students’ social and behavioral outcomes and reduces suspensions and office discipline referrals when implemented with fidelity (Bradshaw et al., 2010, 2012; Gage et al., 2018). Tier 1 supports also translate into significant cost savings for states implementing the model with high fidelity (Bradshaw et al., 2020). Although many states and school divisions have invested considerable resources in scaling up Tier 1 supports, many have struggled to implement Tier 2 supports to scale (Lewis et al., 2023). A critical factor in supporting program adoption, fidelity, and scale-up is the perceived social validity of the model (Hugh et al., 2022; McNeill, 2019). Social validity refers to the extent to which the goals, procedures, and outcomes of interventions are perceived as usable, feasible, and valuable by interested parties (Kazdin, 1977; Wolf, 1978). Wolf (1978) specified social validity as encompassing three features: (a) the social significance of the behavioral goal, (b) the social appropriateness of the procedures, and (c) the importance of the effects on society. Evaluating social validity is essential to an in-depth understanding of the full effects of an intervention; for example, social validity measures may indicate unintended consequences (Hawkins, 1991; Strain et al., 2012). Collecting social validity is also critical to making informed decisions about the development of the intervention and facilitating sustained implementation (Carr et al., 1999; Schwartz & Baer, 1991).
Although collecting social validity data from multiple viewpoints is considered best practice (Schwartz & Baer, 1991; Spear et al., 2013), it may be challenging to do so in practice. When collecting data from multiple stakeholders is in feasible, it may be beneficial for researchers to consider their research questions and then make decisions based on whose input most appropriately addresses those questions (Ledford et al., 2016). Snodgrass and colleagues (2021) recommended considering three aspects when deciding on who should participate in social validity assessments: (a) their connection to the intervention, (b) how knowledgeable they are regarding the intervention, and (c) how their personal characteristics influence their perspective.
To assess multiple stakeholders’ viewpoints meaningfully, researchers need to use valid and reliable social validity measures. Unfortunately, there is no consensus on which types of social validity assessments are the most reliable or valid (Snodgrass et al., 2021). For example, Marchant and colleagues (2012) argued that surveys completed before, during, and after implementation with multiple participants provide researchers with sufficient information to know if the intervention is positively perceived by society. However, Finn and Sladeczek (2001) noted the lack of comprehensive social-validity surveys and urged researchers to move away from surveys as the sole social-validity measure (see also Berger et al., 2016). One approach for obtaining a more in-depth understanding of social validity is qualitative research. Although qualitative methods (e.g., interviews and focus groups) may provide greater depth and detail regarding individuals’ preferences or perceptions (Leko, 2014) and, in turn, inform efforts to make future iterations of the intervention more acceptable and sustainable, it may increase the time and cost associated with data collection.
Regardless of the different viewpoints on how and from whom to collect social validity, social validity has been associated with high-fidelity implementation of programs, like Tier 1 supports, as well as evidence-based practices for individuals with autism (McIntosh et al., 2013; McNeill, 2019). As such, social validity is critical to consider when aiming to scale up school-based models. To better understand potential barriers and facilitators of implementing Tier 2 PBIS supports, the current study systematically reviewed evaluations of the social validity of Tier 2 supports in the context of experimental research. The overarching goal of this study is to provide insights that will inform high-quality implementation and scale-up of Tier 2 PBIS supports and address a pressing research-to-practice gap.
Social Validity of Tier 2 PBIS Interventions
Despite the recognized importance of social validity, it is not always evaluated in experimental research on PBIS or other practices and programs used with students with and at risk for disabilities. For example, in Snodgrass and colleagues’ (2018) evaluation of 429 single-case studies from 2005 to 2016, social validity was systematically evaluated in only 26.8% of studies, and only 24.3% of the articles that evaluated social validity did so comprehensively (i.e., evaluated goals, procedures, and outcomes). In their review of PBIS Tier 2 interventions, Bruhn and colleagues (2014) found that 22 of the 28 studies reviewed and reported social validity outcomes. The authors reported studies that used social validity measures generally included questionnaires, interviews, and rating scales from students, parents, and teachers, and social validity results were generally positive. However, Bruhn et al. did not systematically examine (a) which aspects (i.e., goals, procedures, outcomes) of social validity were measured and reported, (b) who participated in what social validity measurements or (c) the social validity results.
More recently, Blair et al. (2021) conducted a meta-analysis to assess the effect of Tier 2 PBIS interventions and moderators of student behavior outcomes. They found that the intervention (e.g., Check-In/Check-Out [CICO], social skills instruction), implementer (e.g., researcher and teacher), and grade level moderated treatment effects. For example, social skills instruction and group contingency interventions had large effects on problem behavior, whereas CICO demonstrated moderate effects. The authors also reported that 88.5% of included studies used a social validity assessment. However, similar to Bruhn et al. (2014), Blair and colleagues did not specify how or when social validity data were collected, what methods were used, or the results of the social validity assessments.
Overview of the Current Study
Despite (a) the relation of social validity to the adoption and implementation of evidence-based practices (Hugh et al., 2022; McNeill, 2019) and (b) Bruhn et al. (2014) reviewing the prevalence of examining social validity in the Tier 2 research base, a systematic review providing an in-depth evaluation of the social validity of Tier 2 PBIS interventions (e.g., who participated in social validity evaluations, when social validity data were collected, and the results of social validity evaluations) has yet to be conducted. To fill this gap in the literature, we conducted a systematic review to examine the proportion of experimental Tier 2 PBIS studies reporting data on social validity, the elements of social validity evaluated (i.e., goals, procedures, and outcomes), and the results of social validity by type of intervention and element of social validity. Additionally, we were interested in whose perceptions were assessed (e.g., students, teachers, and families), at what stage of the implementation process were social validity data collected, and what type of social validity measures were used. Together, these findings provide important insights to inform future efforts to promote the adoption and scale-up of Tier 2 PBIS supports with fidelity. The current systematic review addresses the following research questions:
Method
We used Maggin and colleagues’ (2017) quality indicators for systematic reviews in special education to guide our review.
Search Procedures
We conducted title, abstract, and keyword searches of the Education Research Complete, Education Full Text, Academic Search Complete, PsycINFO, Teacher Reference Center, Education Resource Information Center, and ProQuest (including dissertations and theses) electronic databases using the following Boolean phrase: (“positive behavior support*” OR PBS OR “positive behavior intervention support*” OR PBIS OR “three-tiered model” OR “multitiered model” OR “multitiered systems of support” OR “MTSS”) AND (“tier 2” OR targeted OR “secondary tier”) AND (support OR intervention OR treatment). We identified 1,137 papers from the initial electronic searches, from which 424 duplicates were removed, for a total of 713 that were screened (see online Supplemental Figure S1 for study flow diagram).
After screening articles from the electronic search, we conducted ancestral (i.e., citations within the included articles) and forward (i.e., articles that cited the included articles) searches on the studies that met inclusion criteria. We also screened the articles included in the Bruhn et al. (2014) and Blair et al. (2021) reviews on Tier 2 interventions that were not identified in the electronic searches. Finally, the first author conducted a hand search of the seven journals with the highest number of identified studies since the Bruhn et al. (2014) review was published.
Screening Procedures
Based on Maggin and colleagues’ (2017) recommendations, the authors independently screened potentially eligible studies using the inclusion and exclusion criteria detailed in Table 1. As the framework for PBIS involves the delivery of universal or Tier 1 procedures before students are identified for additional Tier 2 support (Van Camp et al., 2021), studies that did not explicitly indicate the use of universal or Tier 1 procedures were excluded from the review. Although the importance of social validity was recognized before the publication of Horner and colleagues’ (2005) standards for high-quality single-case design research (e.g., Wolf, 1978), we chose 2005 as the cutoff date for this review because (a) Horner and colleagues’ 2005 paper further established social validity as a critical element of rigorous single-case design and (b) school-wide PBIS was becoming a broadly implemented and researched practice at that time (e.g., Horner et al., 2005; Stormont & Lewis, 2005; Walker et al., 2005).
Inclusion and Exclusion Criteria.
Note. PBIS = Positive behavior interventions and supports; MTSS = multitiered systems of support.
For title and abstract screening, each article was screened independently by two different authors. The first author screened the title and abstracts of each potentially eligible study. The second, third, and fourth authors each screened 33% of the studies. Both screeners must have agreed that the title or abstract demonstrated the publication met the exclusion criterion for exclusion prior to full-text screening. Interrater agreement was calculated using a percentage agreement formula in which total agreements are divided by the total agreements plus disagreements, resulting in 93% agreement. A total of 622 papers were excluded during the title and abstract screening, resulting in a full-text screening of the 91 remaining papers (see online Supplemental Figure S1).
After title and abstract screening, the first author trained the second, third, and fourth authors to conduct full-text screening using: (a) an introduction to the screening protocol; (b) practice full-text screening a subset of five articles identified for full-text screening; and (c) full-text screening of additional articles, if and as needed, until each author reached at least 90% interrater agreement with the lead screener. The lead author screened all eligible publications, and the other three screeners were assigned collectively a random subset of 50% (i.e., each screener screened 17% of the total articles) of the publications. Interrater agreement was 82%. Discrepancies were resolved through discussion. Forty-seven studies were excluded during full-text screening, resulting in inclusion of 44 studies in the review (see online Supplemental Figure S1).
Coding
Following full-text screening, the first author trained the second, third, and fourth authors to code included articles using: (a) an introduction to the coding procedure; (b) practice coding of a subset of three articles identified for inclusion; and (c) coding additional articles, if and as needed, until achieving at least 90% interrater agreement. The first author coded all eligible publications, and the three secondary coders coded 26%, 17%, and 10% of articles, resulting in double-coding of 53% of all studies. Interrater agreement was 92% across coders and items. Discrepancies were resolved through discussion.
For each study we coded: (a) design of the study, (b) number and demographics of the student participants, (c) independent variable, (d) dependent variable, (e) PBIS tier at which the intervention was implemented, (f) treatment fidelity of Tier 1 supports (high, moderate, and low), (g) systematic assessment of social validity (yes and no), and (h) inclusion of a research question related to social validity (yes and no). For participant demographics, we coded age/grade, race/ethnicity, disability identification, and gender. The coding form is available from the first author upon request.
Systematic social validity was defined as an evaluation of social validity described in the “Method” section and findings reported in the results and/or discussion sections. When authors did report the assessment of social validity in their study, we further coded specific aspects of social validity including (a) the element(s) of social validity assessed (i.e., goals, procedures, and outcomes), (b) who participated in the social validity assessment(s) (i.e., interventionists, student participants, family members, and others), (c) when social validity data were collected (i.e., before the intervention, during the intervention, and after the intervention), (d) the type of measurement(s) used, and (e) the results of social validity assessments. When evaluated, we coded the authors’ interpretations of social validity findings as all positive, all neutral, all negative, or mixed for the (a) goals, (b) procedures, (c) feasibility (considered a specific aspect of the social validity of intervention procedures), (d) intent to continue using (also considered a specific aspect of the social validity of intervention procedures), and (e) outcomes of the intervention. We coded overall social validity results as all positive, all neutral, all negative, or mixed if the authors did not distinguish reporting of social validity for goals, procedures, and outcomes.
Results
A description of each of the 44 studies reviewed is provided in online Supplemental Table S1. Five studies were dissertations or theses, and 39 were published in peer-reviewed journals. Forty-three studies were conducted in the United States, and one study (Karhu et al., 2021) was conducted in Finland. Single-case designs were the most frequently used experimental method (n = 32). Regarding group designs, there were nine pre/post-quasi-experimental designs and three randomized control trials. Treatment fidelity of Tier 1 supports was reported in 29 studies; most (n = 23) evaluated fidelity using the Schoolwide Evaluation Tool (SET; Sugai et al., 2001) and most (n = 22) reported high treatment fidelity; see online Supplemental Table S1 for participant demographics.
The most common target behavior was problem behavior, with some authors using the term “challenging” or “disruptive” behavior. Problem behavior was typically defined according to the student’s individual needs; however, authors referred to the behaviors, in general, as problematic or disruptive. Twenty-nine studies used a version of the Behavior Education Program (BEP) or Check-in/Check-Out (CICO) as the intervention, five used a variation of social skills instruction, and four used a variation of a function-based intervention (FBI). Two studies evaluated two interventions: Majeika et al. (2022) assessed CICO in comparison to Breaks are Better, and McDaniel et al. (2018) assessed Coping Power in comparison to CICO (see online Supplemental Table S1). The remaining six studies included PBIS, a “Class Pass Intervention”, “Alternative to Lunch Program for Students”, “Merging Two Worlds”, “Self-Regulated Strategy Development,” and an electronic version of self-monitoring.
RQ 1: Evaluations and Results of Social Validity
We considered evaluation of social validity systematic when described in the method section and findings reported in either the results or discussion section. Twenty-nine studies (66%) systematically evaluated social validity and 12 studies (27%) included a research question regarding social validity. Fifteen (52%) of the studies systematically evaluating social validity reported entirely positive perceptions of social validity, whereas 14 (48%) reported mixed (i.e., positive and negative or neutral) findings for social validity. No study had entirely negative or neutral perceptions from interested parties.
RQ 2: Systematic Evaluations of Social Validity by Intervention Type and Elements of Social Validity
We examined social validity by intervention type and element of social validity (goals, procedures, and outcomes); see online Supplemental Tables S2, S3, and S4.
Findings by Intervention Type
The majority of Tier 2 interventions in the studies reviewed were CICO or BEP (see Figure 1). Of the 29 studies that evaluated CICO or BEP, most (n = 21) systematically evaluated social validity, with 11 (52%) reporting entirely positive and 10 (48%) reporting mixed results regarding social validity. Of the four studies examining a FBI, two (50%) systematically evaluated social validity, and both (100%) reported mixed perceptions. For the other ten Tier 2 studies, six (60%) systematically collected social validity; three (50%) of these studies reported entirely positive perceptions, and three (50%) reported mixed perceptions. The total number of studies reported is higher than those included in the articles because two articles (Majeika et al., 2022; McDaniel et al., 2018) systematically assessed the social validity of two interventions.

Percentage of Studies Reporting Social Validity Findings for Different Elements of Social Validity.
Findings by Element of Social Validity
Most studies reported on the intervention procedures (n = 27; 87%) or the intervention outcomes (n = 16; 52%; see Figure 1). Eleven studies (38% of studies that systematically evaluated social validity) reported the social validity of both procedures and outcomes. Three of the 29 studies (10%) reported all three elements (i.e., goals, procedures, and outcomes). Five studies (16%) did not specify what element of social validity was evaluated or did not distinguish between the results for each element.
RQ 3: Participants, Timing, and Measures in Social Validity Evaluations
Participants engaged in systematic evaluations of social validity included interventionists, teachers who did not implement the intervention, student participants, and family members. Researchers assessed social validity during the planning stage of the study, during the implementation of the intervention, and after the intervention. Researchers used surveys and interviews to evaluate social validity.
Participants
Of the 29 studies that systematically assessed social validity, data were primarily collected from interventionists (n = 24; 77%) and student participants (n = 21; 68%; see Figure 2). This sums to a number higher than the total number of studies that systematically evaluated social validity because 19 studies (66%) collected social validity from multiple groups. For example, Bunch-Crump (2015) collected social validity data from student participants, interventionists, family members, and teachers (see online Supplemental Tables S3 and S4).

Percentage of Studies Reporting Social Validity Findings for Different Groups of Participants.
Three studies (10% of studies that systematically evaluated social validity) collected data from only student participants; one of those studies (33%) reported entirely positive perceptions and two (67%) reported mixed perceptions. Nine studies (29%) only collected social validity from the interventionists; eight of those studies (89%) reported entirely positive perceptions and one (11%) reported mixed perceptions. Ten studies (35%) collected social validity from student participants and nterventionists, with four of those studies (40%) reporting entirely positive perceptions and six (60%) reporting mixed perceptions. One study collected social validity data from family members and interventionists, and one study collected social validity data from student participants and their families; both studies reported entirely positive results. Five studies (16%) collected social validity data from student participants, their families, and interventionist; one of those studies (20%) reported entirely positive perceptions and four (80%) reported mixed perceptions (see online Supplemental Tables S3 and S4).
Timing
Most of the studies that systematically evaluated social validity only did so after intervention (n = 16; 55%; see Figure 3). Six studies evaluated social validity during and after intervention (21%), and two studies evaluated social validity during intervention (7%). Five studies did not report when social validity was collected (17%). No study evaluated social validity before, during, and after data collection (see online Supplemental Table S2).

Percentage of Studies Reporting Social Validity Findings for Different Times Data Were Collected.
Eight of the studies that collected social validity after the intervention had entirely positive results (50%), and the other eight studies had mixed results (50%). Four of the studies that collected social validity during and after implementation had entirely positive results (67%), and two were mixed (33%). All four studies that collected social validity during intervention had mixed results. For the remaining five studies that did not report the timing, four (80%) had entirely positive results, and one (20%) had mixed results.
Measures
Of the 29 studies that systematically collected and analyzed social validity, 27 (93%) used surveys and 2 (7%) conducted interviews. Twenty-one of the 27 studies (78%) that used surveys used pre-existing surveys. The Intervention Rating Profile (IRP-15; Witt & Elliott, 1985) was used five times, with a modified version of that rating scale (Martens et al., 1985) used twice. The BEP Acceptability Questionnaire (Hawken & Horner, 2003) and the Children’s Intervention Rating Profile (Witt & Elliott, 1985) were each used four times. The Behavior Intervention Rating Scale (Elliott & Treuting, 1991) was used twice. The Contextual Fit Questionnaire (Horner et al., 2003), the Treatment Acceptability Rating Form—Revised (Reimers & Wacker, 1988), the Therapeutic Alliance Scale for Children (Shirk & Saiz, 1992), the Usage Rating Profile Intervention—Revised (Chafouleas et al., 2011), and the CICO Program Acceptability Questionnaire (Hawken & Horner, 2003) were each used once. Two of the 27 studies (6%) that used a social validity survey reported they were administered anonymously, whereas 25 studies (93%) did not report how the surveys were administered. One of the two studies using interviews reported some of the interview questions, and the other study did not report any interview questions (see online Supplemental Table S3). Fourteen of the 27 studies (52%) that used surveys reported entirely positive perceptions regarding social validity (see online Supplemental Table S3). One of the interviews (Blair et al., 2010) reported entirely positive perceptions, whereas Gerard (2008) reported mixed perceptions (see online Supplemental Table S4).
Discussion
We examined the social validity of Tier 2 PBIS interventions by evaluating (a) how many studies systematically evaluated and reported social validity, (b) the results of social validity assessments, and (c) how social validity was measured in the research base of experimental Tier 2 PBIS studies (i.e., who participated, which measurements were used, when data were collected; see online Supplemental Table S3). Findings indicated that 66% of experimental studies examining Tier 2 PBIS interventions evaluated social validity, with most studies reporting positive social validity, which may facilitate sustained implementation with positive effects over time.
Previous reviews of social validity (e.g., Snodgrass et al., 2018; Spear et al., 2013) reported that relatively few studies examined social validity. Snodgrass and colleagues (2018) reported that 26.8% of single-case studies published from 2005 to 2016 evaluated social validity, a much smaller proportion than in our review. Our finding that 66% of Tier 2 PBIS studies systematically assessed social validity is consistent with Bruhn and colleagues’ (2014) finding that 79% of Tier 2 PBIS studies reported social validity. A couple of possible reasons may explain why most Tier 2 PBIS studies have evaluated social validity. First, the PBIS framework is designed to include support and input from interested parties (Horner & Sugai, 2015). Thus, researchers may be accustomed to evaluating their perceptions when researching the PBIS framework. Second, several existing surveys assess teacher and student perceptions of CICO, a frequently used Tier 2 intervention. Having existing surveys available may reduce the time and effort needed for researchers to evaluate social validity. Researchers may want to consider using existing, validated surveys to assess social validity efficiently (e.g., User Rating Profile for Supporting Students’ Behavioral Needs; Chafouleas et al., 2018) when conducting applied research.
The generally positive social validity findings are consistent with the results in Bruhn and colleagues’ (2014) review and suggest that interested parties perceive commonly researched Tier 2 PBIS interventions positively, which bodes well for scaling up and sustaining these interventions in practice. It is also possible that the generally positive social validity findings were influenced by selection bias and social desirability bias in the studies reviewed. All interested parties who participated in social validity evaluations had agreed to participate in the studies and, thus, may have been positively disposed toward the intervention. Additionally, as most social validity evaluations were surveys and interviews, respondents may have felt some pressure to respond positively. To reduce potential socially desirable responses, researchers can have participants complete surveys anonymously and have researchers not associated with the study conduct the interviews.
Recommendations for Scaling-Up PBIS
Results suggest that relevant parties held generally positive perceptions of the Tier 2 PBIS interventions used in the studies reviewed (e.g., CICO, BEP, FBI, and social skills instruction). Most included Tier 2 studies (86%) evaluated CICO, BEP, or a variation of these two interventions. Although interested parties had generally positive perceptions of CICO and BEP, not all participants rated the interventions favorably. Some studies reported teachers felt CICO was a lot of work for minimal reward and some students disliked participating. It is possible some teachers and students may respond more positively to different Tier 2 interventions, such as self-management, social skills instructions, and mentoring. For example, McDaniel et al. (2018) found teachers and students generally positively perceived Coping Power, a Tier 2 social-emotional learning curriculum Researchers should continue to evaluate perceptions of social validity for the plethora of Tier 2 interventions. Conducting multiple and in-depth evaluations of the social validity of multiple Tier 2 PBIS interventions will provide researchers and school personnel with important data with which to inform decisions related to the selection of which targeted interventions to implement for different groups of learners.
When collaborating with schools regarding the implementation and scaling up of Tier 2 interventions, PBIS coaches may want to consider including teachers, families, and students in the development of the intervention to potentially improve social validity and, relatedly, initial perceptions of the intervention. For example, Blair et al. (2010) included teaching staff in the participants’ classroom in the design and implementation of the intervention. The authors indicated the teaching staff evaluated the effectiveness, feasibility, and usability of the intervention positively. As teacher buy-in likely influences implementation fidelity (Dariotis et al., 2017; McIntosh et al., 2013; Miramontes et al., 2011), including multiple interested parties in the development of the intervention may result in improved perceptions of the interventions, thereby enhancing fidelity and impact. Interested parties can informally support the development of interventions by participating in planning meetings and providing input regarding the school environment.
Researchers may also want to collect social validity data to systematically gauge interested parties’ perceptions of intervention goals, procedures, and outcomes before implementation. If the results of initial social validity data collection indicate teachers do not positively perceive aspects of the intervention, researchers can make adjustments to improve initial perceptions. Although no study in the current review systematically collected social validity before implementing the intervention, multiple studies indicated they worked with the schools or teachers to develop the intervention. For example, Boyd and Anderson (2013) worked with the school to determine what intervention would be implemented, how and when data would be collected, and trained the teachers to implement the intervention independently. Boyd and Anderson reported that all teachers in the study reported entirely positive perceptions of the intervention. It is possible working with teachers prior to implementation and making adjustments based on their needs improved their perceptions. Additionally, collecting social validity data before, during, and after allows researchers to measure whether and how interested parties’ perceptions of the intervention change over time.
Additionally, when working with schools to scale up Tier 2 PBIS interventions, PBIS coaches may want to highlight evidence of positive social validity for targeted interventions. As many teachers prioritize practice-based evidence (i.e., information from other teachers; Beahm et al., 2021), informing teachers of other teachers’ positive experiences and perceptions of targeted Tier 2 interventions—such as those documented in the studies reviewed—may facilitate buy-in. PBIS coaches may want to prioritize interventions with evidence of high social validity (e.g., CICO and FBI) and provide practitioners with practice-based evidence of positive perceptions, such as quotes from teachers in research articles and from local schools who have implemented targeted Tier 2 PBIS interventions. For example, CICO is an evidence-based practice with relatively high social validity from teachers, families, and students that can engender support and buy-in among these groups.
Recommendations for Future Research
Although the social validity results generally were positive, several studies that used surveys to assess social validity reported a mean rating across items and participants but did not report a measure of variability or disaggregate findings for different areas of social validity or different groups of respondents. This resulted in difficult to interpret social validity findings. For example, a mean rating for survey items of 5.0 (on a 1–7 scale) could indicate all interested respondents felt moderately positively across all areas of social validity, or there might be meaningful differences in perceptions of the intervention between respondent groups and areas of social validity that imprecise reporting obfuscates. More detailed and complete reporting of social validity results can help support the development of future iterations of interventions by highlighting more and less socially valid components in specific areas or for specific groups of respondents. Thus, we recommend reporting both measures of variability and central tendency for quantitative measures of social validity and providing raw data for all respondents and all items in a table, appendix, or supplemental material.
Additionally, researchers may consider using multiple social validity measures (e.g., surveys and interviews) to understand more fully how stakeholders perceive interventions and why. Different measures provide different information regarding the perceptions of interested parties. For example, rating scales gauge how participants perceive interventions, but open-ended questions and interviews can explain why they hold those perceptions. Results from multiple social validity measures can help future researchers improve the feasibility and acceptability of interventions, develop necessary adaptations, inform the development of new interventions based on stakeholder feedback, and ultimately provide educators with more effective and feasible practices. Although most studies that systematically evaluated social validity used a survey, only two conducted an interview and none used both a survey and an interview to evaluate social validity. Interviews can provide valuable information regarding participants’ perceptions of Tier 2 PBIS interventions. For example, Gerad (2007) interviewed student participants to understand their perceptions of BEP. Although some students stated they liked the intervention, others felt they were too far behind to catch up and decided to stop trying because the task felt too daunting. In this example, a goal-setting intervention may not have been enough support to help the students feel successful. These detailed explanations can help researchers understand why some participants are responsive to intervention and others are not. When interviewing participants is not feasible, researchers might include open-ended survey questions on a rating scale to reduce time demands but still allow participants to provide detail and nuance regarding their perceptions of the intervention.
Intervention procedures were evaluated in almost all the studies that systematically assessed social validity. This is an important aspect of social validity because it indicates whether interested parties are confident and comfortable implementing the intervention. Evaluating interested parties’ perceptions of the goals and outcomes provides valuable information as well, although these aspects of social validity were less often evaluated. It is possible, for example, the procedures of a Tier 2 intervention are perceived as usable and feasible, but interested parties do not find the goals (i.e., the target behavior or skill) valuable or the change in behavior (i.e., outcomes) worth the time and effort. For example, Hawken et al. (2011) reported most participants’ behavior improved after the introduction of BEP, but social validity results indicated teachers were mixed regarding the importance of the outcomes achieved. Thus, we recommend interventionists assess the social validity of intervention goals, procedures, and outcomes—not just one or two dimensions of social validity.
Limitations
There are several important limitations to consider when interpreting the results of this review. First, we reviewed only experimental quantitative studies. It is possible including non-experimental and qualitative studies of Tier 2 PBIS interventions would have changed our findings. Secondly, although we used multiple search engines and search methods to identify studies that met our inclusion criteria, it is possible that we did not identify some relevant Tier 2 PBIS interventions. Third, we coded systematic social validity as occurring only if the manuscript explicitly referred to “social validity” or “treatment acceptability.” Several studies reported data that might be considered to examine social validity, but we did not code it as such because the authors did not identify it in this way. For example, Fairbanks et al. (2007) compared participants’ performance to typically developing peers in the same class. This comparison could be perceived as examining social validity related to the change in outcomes. However, because the authors did not indicate that these data evaluated social validity, we did not code it as such. Including such data may have changed our findings. Finally, although most studies reported positive social validity, the studies varied in how they reported and interpreted the social validity findings. We coded social validity findings as positive, negative, neutral, or mixed based on the authors’ interpretations. Coding social validity findings based on independent criteria may have changed our findings and conclusions.
Conclusion
This systematic literature review evaluated how social validity was reported in Tier 2 PBIS experimental studies. Results of this study indicate that 66% of studies systematically evaluated social validity, and 52% of those studies reported largely positive perceptions of Tier 2 PBIS interventions from teachers, families, and students. These positive social validity perceptions may favorably influence the adoption and scale-up of these interventions in schools. Program developers and implementers should consider collecting data on social validity at multiple points during the implementation, using multiple measures (e.g., surveys and interviews), and collecting the perceptions of multiple relevant parties (e.g., student participants and teachers) to more robustly evaluate the social validity of Tier 2 PBIS interventions.
Supplemental Material
sj-docx-1-pbi-10.1177_10983007241230596 – Supplemental material for A Systematic Review of the Evaluation of Social Validity in Experimental Examinations of Tier 2 Schoolwide Positive Behavior Interventions
Supplemental material, sj-docx-1-pbi-10.1177_10983007241230596 for A Systematic Review of the Evaluation of Social Validity in Experimental Examinations of Tier 2 Schoolwide Positive Behavior Interventions by Lydia A. Beahm, Bryan G. Cook, Alan McLucas, Kaci Ellis and Catherine P. Bradshaw in Journal of Positive Behavior Interventions
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
This work was supported by funding from the University of Virginia’s School of Education and Human Development Innovative, Developmental, Exploratory Awards to the lead author.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
