Abstract
Full and durable implementation of school-based interventions is supported by regular evaluation of fidelity of implementation. Multiple assessments have been developed to evaluate the extent to which schools are applying the core features of school-wide positive behavioral interventions and supports (SWPBIS). The SWPBIS Tiered Fidelity Inventory (TFI) was developed to be used as an initial assessment to determine the extent to which a school is using (or needs) SWPBIS, a measure of SWPBIS fidelity of implementation at all three tiers of support, and a tool to guide action planning for further implementation efforts. In this research, we evaluated the psychometric properties of the TFI in three studies: a content validity study, a usability and reliability study, and a large-scale validation study. Results showed strong construct validity for assessing fidelity at all three tiers, strong interrater and 2-week test–retest reliability, high usability for action planning, and strong relations with existing SWPBIS fidelity measures. Implications for accurate evaluation planning are discussed.
Keywords
Schools across the country are facing the demand to provide rigorous educational opportunities to a highly diverse population of learners requiring various levels of academic and behavior support. The most recent reauthorization of the Individuals With Disabilities Education Act (2004) provided the impetus for an increased focus on empirically supported practices. However, simply electing to adopt evidence-based practices without attending to the implementation process is unlikely to improve outcomes (Fixsen, Blase, Duda, Naoom, & Van Dyke, 2010). Implementation abandonment, wherein schools discontinue the use of effectively implemented practices in place of new ones each year, is commonplace in schools across the country (Adelman & Taylor, 2003). This phenomenon carries costs with regard to system resources, including financial losses and reduced staff buy-in, as well as student outcomes (McIntosh et al., 2013). Empirical research shows that assessing fidelity and using those data to inform action planning can increase sustainability and decrease the likelihood of abandoning effective practices (McIntosh, Kim, Mercer, Strickland-Cohen, & Horner, 2015).
One effective and widely implemented practice is school-wide positive behavioral interventions and supports (SWPBIS; Sugai & Horner, 2009), a three-tiered framework that promotes the use of positive and preventive approaches to behavior support at a systems level. More than 21,000 schools in the United States have adopted SWPBIS in efforts to establish positive, safe, predictable, and consistent school climates (Horner, 2014). Research indicates that high fidelity of implementation of SWPBIS is associated with improved student and teacher outcomes, including an increase in student perception of school safety, a reduction in number of office discipline referrals (ODRs), a decrease in student use of school counseling services, growth in academic achievement, and an increase in teacher self-efficacy (Bradshaw, Mitchell, & Leaf, 2010; Horner et al., 2009; Kelm & McIntosh, 2012; McIntosh, Bennett, & Price, 2011; Nelson, Martella, & Marchand-Martella, 2002; Ross, Romer, & Horner, 2012). Flannery, Fenning, Kato, and McIntosh (2014) found that SWPBIS reduced the level of problem behavior in high school students, and the level of reduction was significantly related to fidelity of implementation, as schools with higher fidelity had decreased rates of problem behavior.
Measuring Fidelity of Implementation of SWPBIS
One of the defining activities of SWPBIS is the use of data for decision making (Algozzine et al., 2010). Data are used to guide both decisions focused on improving student supports and decisions about how best to implement SWPBIS features. For schools to implement SWPBIS successfully, ongoing evaluation of fidelity of implementation and informed action planning based on data are essential. Fidelity of implementation is defined as the extent to which a program, intervention, framework, or practice, “as conceptualized in a theoretical model or manual, is implemented as intended” (Schulte, Easton, & Parker, 2009, p. 460). Although the importance of fidelity is not a new concept in educational research (O’Donnell, 2008), school-based assessment of implementation has recently become the subject of increased focus. The trend of assessing fidelity of school systems is reflected in the rapid increase in the number of assessment tools available for evaluating the core components of SWPBIS implementation. These fidelity measures include (a) the School-Wide Evaluation Tool (SET; Sugai, Lewis-Palmer, Todd, & Horner, 2001), (b) the School-Wide Benchmarks of Quality (BoQ; Kincaid, Childs, & George, 2005), (c) the Team Implementation Checklist (TIC; Sugai, Horner, & Lewis-Palmer, 2001), (d) the PBIS Self-Assessment Survey (SAS; Sugai, Horner, & Todd, 2000), (e) the Benchmarks for Advanced Tiers (BAT; Anderson et al., 2012), (f) the Individual Student Systems Evaluation Tool (ISSET; Lewis-Palmer, Todd, Horner, Sugai, & Sampson, 2003), and (g) the Monitoring Advanced Tiers Tool (MATT; Horner, Sampson, Anderson, Todd, & Eliason, 2013). Collectively, these measures assess implementation at each of the three tiers of SWPBIS, but there has not been a single measure that can be used to assess fidelity of implementation of all three tiers on the same scale, which has presented challenges for evaluation across schools at the district, regional, or state level.
Tiered Fidelity Inventory (TFI)
The SWPBIS TFI (Algozzine et al., 2014) was developed to be a comprehensive fidelity of implementation tool to be used alone or in conjunction with other SWPBIS assessments. Although the existing fidelity measures can be used to assess fidelity of implementation of SWPBIS, there was no single tool that school teams could use to measure initial implementation, develop an action plan, and monitor implementation progress across all three tiers. The TFI was designed to be a more comprehensive and efficient measure of fidelity, with a common format, scale, and language to assess each tier, for schools at any level of implementation. The TFI is intended as (a) an initial assessment to determine whether a school is using (or needs) SWPBIS, (b) a guide for implementation of Tier I, Tier II, and Tier III practices, or (c) an index of sustained SWPBIS implementation. The TFI was compiled from existing SWPBIS fidelity measures and unpublished fidelity measures used in Florida, Illinois, Maryland, Missouri, and North Carolina. Table 1 provides a description of the most commonly used existing SWPBIS fidelity measures, along with the TFI. As with these other tools, the TFI is freely available for download at http://www.pbis.org.
SWPBIS Fidelity Measures.
Note. SWPBIS = school-wide positive behavioral interventions and supports; Foundations = systems-level components needed for implementing at Tiers II and III; PBIS = Positive Behavioral Interventions and Supports.
Systems-level components needed for implementing at Tiers II and III.
The TFI is organized into three scales, representing Tier I (universal), Tier II (targeted), and Tier III (intensive). Each scale can be assessed separately or together to evaluate overall implementation at all three tiers. These options allow for various intended uses: (a) as a complete index of all tiers to establish implementation status and determine focus, (b) as a quarterly progress monitoring tool to guide action planning for implementation of tiers of current focus, and (c) as an annual formative evaluation for tiers already in place. Teams use a Likert-type scale and detailed rubric to indicate whether the content of each item is not implemented, partially implemented, or fully implemented. Data sources are included to help teams evaluate each item objectively. Tier I (universal SWPBIS features) assesses 15 critical features of school-wide supports such as “School has five or fewer positively stated behavioral expectations and examples by setting/location for student and staff behaviors (i.e., school teaching matrix) defined and in place.” Subscales in the Tier I scale include Teams (two items), Implementation (10 items), and Evaluation (three items). Tier II (targeted SWPBIS features) evaluates 13 core features of targeted interventions such as “Tier II team uses decision rules and multiple sources of data (e.g., ODRs, academic progress, screening tools, attendance, teacher/family/student nominations) to identify students who require Tier II supports.” Subscales in the Tier II scale include Teams (four items), Interventions (five items), and Evaluation (four items). Tier III (intensive SWPBIS features) includes 17 items (e.g., “A written process is followed for teaching all relevant staff about basic behavioral theory, function of behavior, and function-based intervention”). Subscales in the Tier III scale include Teams (four items), Resources (three items), Plans (six items), and Evaluation (four items).
Because research has shown that self-assessment of fidelity can be artificially inflated (Noell et al., 2005; Wickstrom, Jones, LaFleur, & Witt, 1996), it is important to ensure that results from fidelity measures are accurate; otherwise, decisions will be flawed. The TFI is intended for use by school teams with the support of an external SWPBIS coach, who facilitates the administration, ensures accuracy of scoring, and guides the team through interpreting the results. Due to varying team membership, the group assessing Tier I supports may be different from the assessors of Tier II and Tier III supports. The TFI can be completed online (http://www.pbisapps.org) or using pencil and paper. After a complete administration of the TFI, summary scores for each scale are provided, representing the percentage of critical features implemented at Tiers I, II and III, as well as a total score for all three tiers. Subscale and item reports are generated to guide coaching and action planning for school teams.
Purpose of the Technical Adequacy Studies
To assess the reliability and validity of the TFI to measure implementation at all three tiers and continue to refine it based on results, we evaluated the psychometric properties of the measure through three studies: (a) a content validity study, (b) a usability and reliability study, and (c) a large-scale validation study. First, an expert panel evaluated the content validity of the TFI, including evaluating the importance of each specific item, how it related to a particular aspect of fidelity, and the usefulness and appropriateness of scoring. Second, the TFI was pilot tested with a small group of school teams and coaches to evaluate the usability of the measure as well as calculate the interrater and test–retest reliability of the tool. Third, the TFI was released nationally for administration under typical conditions to assess its relation to existing SWPBIS fidelity measures. The remainder of the article describes the methods and results of these studies. Because these studies used different samples and methodologies, they are described separately.
Content Validity Study
Method
Participants
Twelve experts in SWPBIS implementation were invited to participate in the content validity study and assess how each item was related to implementation, as well as rate the measure as a whole. Participants had to be one or both of the following: (a) a researcher in SWPBIS with at least two published studies using and reporting SWPBIS fidelity of implementation data in the past 10 years (n = 5) or (b) an experienced SWPBIS implementer with at least 15 years of experience as a school- and district-level implementer and team trainer (n = 7). Individuals were not eligible to participate if they assisted in developing the TFI or shared an institutional affiliation with any developers. There was a 100% response rate, with 2% of responses with missing data.
Measure
We used a survey to assess content validity, the extent to which the specific items of the TFI adequately represent implementation of SWPBIS, which assists in assessing whether the items should be retained, revised, or removed, as well as whether the measure as a whole is valid (Polit & Beck, 2006; Waltz, Strickland, & Lenz, 2005). The survey (based on previous content validity research; McIntosh, MacKay, et al., 2011) included three sections. For each item, we asked (a) the extent to which it addressed important aspects of fidelity of implementation (to assess item validity), (b) the extent to which it was related to the proposed subscale (to indicate factor structure), and (c) the extent to which the scoring criteria were valid (to assess validity of scoring). For each scale, we asked the extent to which the items assessed important aspects of the tier and whether any items should be added or removed. For the measure as a whole, we asked six overall questions (e.g., directions, response format, overall content validity). We used a 4-point Likert-type scale (strongly disagree, disagree, agree, strongly agree) for each question and also asked for open-ended feedback, such as suggestions for rewording items and specifying items to add or remove from the measure.
Procedure
We invited participants to complete the survey anonymously through a secure online surveying program. Two separate analyses were conducted to evaluate the data from the content validity survey. First, interrater agreement (IRA) was calculated to determine the extent to which the experts’ ratings were consistent. As recommended when the number of expert panel participants is 5 or more (Davis, 1992; Lynn, 1986), the 4-point scale was dichotomized by combining strongly disagree and disagree as one rating and agree and strongly agree as one rating. The IRA was calculated for each item and for the survey as a whole. Next, a Content Validity Index (CVI) score was calculated for each item based on the representativeness of the assessment tool. The number of experts who rated an item as agree or strongly agree was counted for each item. This sum was divided by the total number of experts to derive the CVI for each item. The overall CVI for the instrument was determined by averaging the CVI for each item. A CVI of .80 or higher is recommended in the literature for new assessment measures, and items below .80 should be examined for revision (Davis, 1992).
Results
Overall, the expert panel reliability (i.e., the extent to which the raters agreed on their ratings) was 93% (Tier 1 = 95%, Tier II = 93%, Tier III = 91%), with 95% of items above the .80 standard. Furthermore, the reliability was 96% for item validity, 95% for factor structure, and 89% for scoring. These figures indicate a high level of agreement among the experts regarding the TFI and its items. The overall CVI was .92, with 95% of questions rated above the criterion of .80 (range = .67–1). The mean CVI for Tier I items was .95 (range = .67–1). Of the two Tier I items rated below the CVI criterion, one was rated as not aligned to the critical features of implementation, and one was rated as unclear in wording. The mean CVI for Tier II items was .93 (range = .75–1). One Tier II item was rated below the criterion. The scoring criteria for this item were noted as unclear. Finally, the mean CVI for Tier III items was .91 (range = .67–1). Three items were scored below the criterion, and feedback from the expert panel indicated the need for more universal language related to intensive interventions (e.g., person-centered planning, Rehabilitation for Empowerment, Natural Supports, Education, and Work [RENEW], wraparound services). Overall, the content validity data demonstrate that the expert panel considered the items, scoring criteria, and overall structure to be a valid measure of the important aspects of fidelity of implementation of SWPBIS.
Changes to Measure
All six TFI items that were rated below the .80 content validity criterion were changed. Based on the feedback from experts, one item was removed from the measure, one item description was revised, scoring criteria for one item were changed, and three items were reworded in both the description and scoring criteria. These items were revised to reflect a common, universal language related to interventions, and scoring criteria were revised to align with the item description. Along with these changes, an item assessing meeting procedures was added to all three tiers and an item evaluating a range of Tier II interventions was included. These additions were based on the open-ended feedback. All changes were made prior to pilot testing.
Usability and Reliability Study
Method
Participants
This study included school teams and their external coaches from 15 schools in five districts across five states (Connecticut, Michigan, Missouri, North Carolina, and Oregon). School SWPBIS teams were recruited by their state SWPBIS leadership teams to provide a range of implementation (i.e., from first year of implementation of Tier I SWPBIS to strong implementation at all three tiers; mean years implementing = 5.56). Schools included elementary (n = 6), K-8 (n = 2), middle (n = 1), junior high/high (n = 4), high (n = 1), and K-12 schools (n = 1). Enrollment for schools with National Center for Education Statistics (NCES) data (n = 14) ranged from 33 to 1,586 (M = 511.79), and percentage of students eligible for free and reduced-price lunch ranged from 5% to 91% (M = 55.79%). Each school team completed the TFI and a usability survey, although in some schools, separate teams completed each scale (i.e., Tier I team completed the Tier I scale, and the Tier II/III team completed the others).
Measure
We developed a usability survey to assess the extent to which the process of administering, scoring, and interpreting the TFI was easy and straightforward. It included 14 questions with a 4-point Likert-type scale (from strongly disagree to strongly agree). For each scale, school teams reported completion time, the extent to which the items assessed important aspects of implementation, and whether items should be added or removed. We also asked them to provide open-ended feedback to improve the measure. The internal consistency of the usability survey (in terms of coefficient alpha) was .87, indicating acceptable reliability. There were no missing usability or TFI data.
Procedure
Pilot study participants completed the TFI and usability survey immediately afterward. The usability and reliability of the TFI was determined through multiple methods of evaluation: (a) the usability survey, providing both quantitative and descriptive data; (b) one TFI completed by the coach prior to using it with the team; and (c) two administrations of the TFI by the coaches facilitating the school teams, provided exactly 2 weeks apart.
Three different analyses were conducted: (a) usability interpretation, (b) calculation of interrater reliability, and (c) calculation of test–retest reliability. Usability encompasses the effectiveness, efficiency, and user satisfaction of a measure (Frøkjær, Hertzum, & Hornbæk, 2000). For consistency with the content validity analyses, we dichotomized the 4-point survey scale and calculated the percentage of responses that were coded as disagree or agree. Items with less than 80% agreement were reevaluated, with changes to the items as needed. We calculated interrater reliability, the extent to which different raters are consistent when using the measure (James, Demaree, & Wolf, 1984; Shrout & Fleiss, 1979), by comparing the score of the coach’s independent TFI (i.e., before meeting with the team) with the score of the administration with the coach leading the team. To do so, we used a two-way random consistency intraclass correlation (ICC) analysis in SPSS. Finally, we calculated test–retest reliability, the extent to which scores vary when the measure (i.e., TFI) is used across time, by comparing the scores of the team’s initial TFI results with those of the 2-week retest. We calculated these ICCs using a two-way random consistency analysis in SPSS.
Results
Usability
Average completion time for each scale was under 15 min (Tier I: 14.5 min, Tier II: 11 min, Tier III: 12.5 min). Responses assessing the overall TFI measure showed strong usability (easy and straightforward process: 100% agree, easy and straightforward scoring: 93% agree, validity for assessing fidelity: 100% agree). Out of 14 questions assessing usability, two had less than 80% agreement (range = .67–1). These questions evaluated the extent to which participants rated that an item should be removed from the TFI. Four participants suggested that an item should be removed from Tier II, and three participants suggested that an item should be removed from Tier III. The most common open-ended feedback from the usability survey was that the TFI was easy to use, and respondents appreciated that they could use one measure to assess fidelity at all three tiers. Respondents were split as to whether the TFI could replace existing fidelity measures. Many noted that it could replace existing Tier II and III measures, but they reported that other Tier I measures could be used for the specialized purposes noted in Table 1 (e.g., TIC for initial implementation, BoQ for deep implementation, SAS for obtaining perceptions from whole staff).
Interrater reliability
The ICCs for interrater reliability across all raters, all tiers, and all items (Tier I, Tier II, Tier III, and overall) were all .99. These scores indicate high reliability in scores between coaches (when completing the TFI about a school alone) and the teams (when assessing fidelity with the TFI with the coach as facilitator).
Test–retest reliability
The ICC for test–retest reliability was .99. These test–retest reliability scores indicate very strong agreement across administrations of the TFI over time, which indicates that the construct is being measured consistently by the measure.
Changes to Measure
Based on the information in the usability survey, TFI items were reworded for clarity in the item description, scoring, or data sources. The majority of changes were clarifying terminology (e.g., person-centered planning, wraparound) and aligning the item descriptions and scoring criteria. One item was added to the Tier I section to split the stakeholder involvement item into two separate items, one item measuring faculty involvement and another measuring student, families, and community member involvement.
Large-Scale Validation Study
Method
Participants
The pilot study included 789 schools across seven states, primarily in Florida and Illinois, in the 2013–2014 school year. Each school completed the TFI, along with at least one of four other fidelity of implementation measures (e.g., BoQ, SAS, TIC, and BAT). Scores from the usability and reliability study (the first administration with coach and team) were also included in analyses. Table 2 provides descriptive statistics for this sample.
School Characteristics for Validation Study Sample (n = 789).
Note. Years implementing SWPBIS available for 96% of schools (n = 759). School demographic data obtained from National Center for Education Statistics for 91% of schools (n = 717). SWPBIS = school-wide positive behavioral interventions and supports; FRL = free and reduced-price lunch; TFI = Tiered Fidelity Inventory.
Measures
Four research-validated measures were used as concurrent measures of SWPBIS implementation: (a) the School-Wide BoQ (Kincaid et al., 2005), (b) the TIC (Sugai, Horner, et al., 2001), (c) the SAS (Sugai et al., 2000), and (d) the BAT (Anderson et al., 2012). The BoQ, SAS, and TIC were used as comparisons for the Tier I scale of the TFI. The BAT Tier II and Tier III scale scores were used as comparisons with the TFI Tier II and III scales. The overall BAT score, which includes the Foundations, Tier II, and Tier III subscales, was compared with the TFI total score (i.e., Tiers I, II, and III).
BoQ
The BoQ is a 53-item Tier I SWPBIS fidelity of implementation scale. The psychometric properties of the BoQ indicate the tool is reliable and valid for measuring Tier I SWPBIS fidelity, with interrater and test–retest reliability above 90% and moderate correlations with the SET (Sugai, Lewis-Palmer, et al., 2001), another Tier I measure (R. Cohen, Kincaid, & Childs, 2007). A total of 321 schools in the sample completed both the BoQ and TFI.
SAS
The SAS is a 43-item self-assessment measure of SWPBIS implementation. For these analyses, the 18-item School-Wide Systems scale was used to assess Tier I implementation. The SAS has high internal consistency and correlations with other validated SWPBIS fidelity measures (Hagan-Burke et al., 2005; Safran, 2006). Internal consistency for all tiers is high (α = .85), and subscale scores range from moderate to high (α range = .60−.92). Concurrent validity with Tier I SET is moderately high (r = .75). A total of 559 schools in the sample completed both the SAS and TFI.
TIC
The TIC is a 17-item measure of Tier I SWPBIS implementation. It assesses the extent to which key start-up activities are implemented. The TIC is intended for use as a progress monitoring assessment measure, and a score of 80% or higher indicates implementation of SWPBIS to criterion levels. Internal consistency for the TIC is high across studies (α range = .91−.94; McIntosh, Mercer, Nese, Strickland-Cohen, & Hoselton, in press; Tobin, Vincent, Horner, Dickey, & May, 2012), and a recent confirmatory factor analysis showed a strong factor structure (McIntosh et al., in press). A total of 164 schools completed both the TIC and TFI.
BAT
The BAT is a 112-item fidelity of implementation measure that assesses implementation at Tiers II and III, as well as foundational structures for supporting systems at Tiers II and III. As with the TFI, each tier can be completed separately, if desired. No published technical adequacy data are available for the BAT. A total of 198 schools completed both the BAT and TFI.
Procedure
School teams and external SWPBIS coaches in two states (Florida and Illinois) were provided with access to the TFI as an additional fidelity of implementation measure in addition to the existing fidelity measures that they were already using. Training for TFI administration was not tightly controlled—Participants were provided access to the measure and a webinar, with no requirement of training or contact with the study authors. When completing the TFI, respondents indicated whether the measure was completed by the school team with an external coach (n = 437) or by the school team alone (n = 282).
Data analysis
Analyses in this study assessed multiple elements of reliability and validity in assessing SWPBIS fidelity. Analyses produced information regarding (a) internal consistency (through coefficient alpha),and (b) concurrent validity with existing measures of SWPBIS implementation (through Pearson correlations). There were no missing TFI data.
Results
Internal consistency
Coefficient alpha was used to evaluate the internal consistency of the measure. The overall internal consistency of the measure was .96. Alphas for Tiers I, II, and III were .87, .96, and .98, respectively, providing evidence of strong internal consistency.
Correlations
Pearson correlations were calculated between the TFI and other existing measures of fidelity of implementation. Correlations were calculated separately by administration condition (i.e., team without external coach and team with external coach). Results are summarized in Table 3. All correlations between the TFI and other measures were statistically significant and were stronger when the team completed the TFI with an external coach. According to criteria from J. Cohen (1988), correlations were generally moderate without a coach, and all were strong with a coach. Furthermore, teams consistently rated their implementation as higher when they completed the measure without an external coach than when they completed an administration with an external coach, indicating a small degree of self-inflation.
Correlations Between TFI and Existing Measures of Fidelity of Implementation by Administration Condition.
Note. TFI = Tiered Fidelity Inventory; BoQ = Benchmarks of Quality; SAS = Self-Assessment Survey; TIC = Team Implementation Checklist; BAT = Benchmarks for Advanced Tiers.
p < .05. **p < .01. ***p < .001.
Discussion
Research has demonstrated that schools with higher SWPBIS fidelity scores have better student outcomes (e.g., lower rates of problem behavior, higher achievement, higher emotional regulation; Bradshaw, Waasdorp, & Leaf, 2012; Childs, Kincaid, & George, 2010; Flannery et al., 2014; Horner et al., 2009). Without reliable and valid assessment of fidelity, there is a danger of assuming that implementation is adequate when it is not. The purpose of this study was to validate and refine a new, comprehensive measure of fidelity of implementation of SWPBIS, the SWPBIS TFI. The TFI was intended to serve as a single measure for assessing SWPBIS implementation at all three tiers, which could provide advantages in terms of efficiency and ease of evaluation for districts and states. Three separate studies were conducted to assess the measure’s construct validity, usability, reliability, and concurrent validity with existing, validated measures of SWPBIS fidelity of implementation. After each study, the measure was refined to continue to enhance its technical adequacy. Collectively, results showed that the measure can be used reliably and validly to assess SWPBIS fidelity of implementation. Results are described by reliability, validity, and usability.
Psychometric Properties of the TFI
Reliability
Educators and administrators need to have confidence that their selected fidelity measures will produce similar scores across conditions. Evidence for reliability comes from the usability and reliability study and the large-scale validation study. The usability and reliability study provided evidence of both IRA (between the coach alone and team facilitated with coach) and test–retest reliability (the team’s ratings over time). Finally, the internal consistency of the measure (from the large-scale validation study) demonstrated high internal consistency overall and within individual tiers. These results provide evidence that the TFI can provide consistent results across raters and time.
Validity
Multiple aspects of validity were assessed. Content validity results (from the expert panel ratings) indicated that the items, scoring criteria, and perceived factor structure of the TFI are valid for assessing the construct of SWPBIS implementation. Concurrent validity analyses (comparisons between the TFI and the BoQ, TIC, SAS, and BAT) showed statistically significant correlations with the other existing SWPBIS fidelity measures, providing indications that the TFI is a valid measure of SWPBIS fidelity.
In line with previous research, relations with other measures were stronger when school teams completed the measure with the guidance of an external coach. Completing the measure without a coach produced adequately valid scores, but the scores appear to have been somewhat inflated, as seen through slightly higher mean scores and lower correlations with other measures. As a result, scores from the TFI appear to be most valid when it is completed with an external coach.
Usability
Although reliability and validity are important, a measure’s utility for decision making is a key factor for applied measures. Evidence for the TFI’s usability came primarily from the usability and reliability study. Users reported that the TFI was easy and straightforward to complete and score, and that it assessed important aspects of fidelity at all three tiers. Descriptive feedback indicated that the TFI was efficient and useful for decision making and action planning to improve systems. Such results indicate that the TFI would be useful for its intended purposes.
Limitations and Future Research
Some limitations of the three studies are apparent. For example, participants in the usability and reliability study were likely to be enthusiastic. It is possible that such selection, although it may have increased the quantity and quality of descriptive feedback, may have biased the results. In addition, the authors themselves did not conduct any external evaluations of SWPBIS fidelity. As a result, the teams completing the TFI may have been the exact same groups participating in administration of the other fidelity measures. In regard to these measures, the lack of detailed technical adequacy data for the BAT makes our findings regarding the TFI Tier II and III scales more tentative than for Tier I. Furthermore, the usability and reliability study’s interrater reliability assessment was conducted with a coach as part of both administrations. Although it is difficult to identify another way to evaluate interrater reliability for a team-based assessment, it is possible that the coach’s presence in both administrations inflated the interrater reliability estimates. Finally, the time of year for concurrent validity was not controlled. As a result, the other measures may have been completed close or far away in time from the TFI administration.
Although these results are promising, further validation work would be useful to assess the technical adequacy of the TFI. First, it will be necessary to validate the finalized TFI measure based on the slight changes to the measure from the final round of feedback. Second, the criterion for adequate implementation (e.g., 70% of total points) has not yet been studied. It will be necessary to identify empirical criteria for adequate implementation. In absence of this research, 70% appears to be a reasonable criterion for adequate implementation at each tier, although mean implementation at Tiers II and III was considerably lower. Third, a rigorous, quantitative assessment of the TFI’s factor structure is necessary (and currently underway). Fourth, it would be useful to further examine the role of coaches in facilitating accurate assessment of fidelity and what factors enhance accuracy in self-rating of fidelity.
Implications for Practice
These results provide indications that the TFI has strong technical adequacy for measuring SWPBIS fidelity at all three tiers and is an appropriate index of implementation. Coaches and coordinators at the school, district, regional, and state levels should feel confident in the measure’s properties and the accuracy of its results. The measure can be used to produce valid results for total, tier, and subscale scores in typical administration (i.e., without extensive training and support in administration). However, the validation study results confirm the TFI authors’ recommendations that administration be conducted with an external coach, due to the objectivity of an outside evaluator. When teams lack an external support to provide additional perspective, the phenomenon of “self-inflation” of fidelity appears to be more likely.
SWPBIS leaders at the school, district, and state can consider whether the TFI can supplement or replace current SWPBIS fidelity measures required for their evaluation plans. Respondents reported that they appreciated the TFI’s comprehensive (i.e., all three tiers in one measure) nature, but that some existing Tier I measures would remain useful for school teams, depending on their specific needs at the time. All of these measures will remain available for administration, scoring, and reporting at http://www.pbisapps.org.
Footnotes
Acknowledgements
The authors wish to thank Stephanie Austin, Linda Bradley, Karen Childs, Bridget Drobac, Susannah Everett, Sarah Moore, Jennifer Rollenhagen, and Erin White for their assistance in data collection.
Authors’ Note
The opinions expressed are those of the authors and do not represent views of the Office or U.S. Department of Education.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by the Office of Special Education Programs, U.S. Department of Education (H326S130004).
