Abstract
The effects of self-monitoring (SM) on teacher behavior are well documented, but previous research does not attempt to control for reactivity as a threat to internal validity. This study examined the effects of a multicomponent SM intervention on the use of a classroom management practice with participant masking to address this absence in the literature. Participating teachers selected between two practices (behavior-specific praise and opportunities to respond). A multiple baseline design across four pre-service teacher interns occurred in general education classroom settings. Participant masking to the purpose of the study precluded exposure to SM, performance feedback, and goal setting. Analyses included an independent visual analysis by three masked raters, an independent quality review for What Works Clearinghouse standards, a nonparametric statistical analysis based on data characteristics, and correspondence reporting between visual and statistical analyses. Overall results indicated an increase in the rate of classroom management practice use by the participants and good social validity across the three constructs. Student outcome data for on-task behavior were inconclusive. Limitations and implications for research and practice are discussed.
Prior research in medicine and education demonstrates self-monitoring (SM) as a promising intervention related to increased awareness of personal performance and desired behavior changes (Caplin & Creer, 2001; Donaldson & Normand, 2009; Guzman et al., 2018; McDougall et al., 2017; Rispoli et al., 2017). SM requires the individual to consciously observe their behavior, record the results, and use the data to make informed decisions to improve outcomes in the future (Bruhn et al., 2015; Rispoli et al., 2017). The process of SM originates in the self-regulation literature (Bandura, 1969, 1977), and although there is disagreement on the mechanism of action responsible for treatment effects (e.g., cognition [Mace & Kratochwill, 1985] or reactivity [Kazdin, 1979]), both operant and cognitive-behavioral models indicate that the change in behavior occurs and is due, at least in part, to the individual’s evaluation of their behavior (Davis, 2015; Mace & Kratochwill, 1985). When an individual is asked to observe and record a specific behavior, the very act of observing and recording influences the behavior of interest (Kafner, 1970). SM relates to self-management that is inclusive of goal setting (GS), monitoring, evaluation, and recording that an individual does on themselves. Although some have argued that all are required for meeting the operational definition, others have indicated any one or more could qualify an intervention as self-management (Briesch & Chafouleas, 2009; Harrison et al., 2019).
Despite some unknowns about why the process works, SM is a well-documented and effective intervention to change teacher behavior with classroom management practices (Rispoli et al., 2017). Participants with reported positive outcomes have included pre-service teachers (Alexander et al., 2012; Hager, 2012; Keller et al., 2005), novice teachers (Briere et al., 2015; Kalis et al., 2007), middle school teachers (Simonsen et al., 2013), and teaching assistants (Pinkelman & Horner, 2017; Seligson-Petscher & Bailey, 2006). Students’ outcomes have been less commonly reported, but have ranged broadly from high school students with emotional and behavioral disorders (Kalis et al., 2007) to students with developmental disabilities in elementary self-contained classrooms (Keller et al., 2005).
SM is often a component in a package of interventions to change behavior. SM treatments frequently include GS, performance feedback (PF), reinforcement, and technology (Bruhn et al., 2015). GS and PF are common within practitioner studies related to classroom management (Briere et al., 2015; Kalis et al., 2007; Oliver et al., 2015; Pinter et al., 2015; Sutherland & Wehby, 2001). GS, like SM, has demonstrated lengthy empirical support and has been referred to as “one of the building blocks for designing effective behavior change interventions” (Epton et al., 2017, p. 1189). Positive effects for GS have been demonstrated across behaviors, populations, and settings (Bruhn et al., 2016; Liao et al., 2019). GS combines easily with SM due to parallels between interventions (targeted behavior change, observing, recording, reporting, and evaluation). Another commonly combined treatment with SM is PF. This is an evidence-based practice for improving treatment integrity and a promising practice for increasing the use of teacher-specific behaviors in pre-service teachers (Cornelius & Nagro, 2014). PF refers to an observer’s critique of an individual’s progress toward a desired outcome. Prior research indicates although SM alone produced behavior change, SM + PF was necessary to achieve criterion levels in a study with three of the four novice teachers (Mouzakitis et al., 2015). SM, GS, and PF have demonstrated positive outcomes for changing behavior in the research literature and share several procedural steps, lending themselves to a natural integration in classroom behavior change treatment packages.
When investigated for research purposes, SM typically has involved a participant recording a target behavior with an external observer simultaneously recording the behavior for reliability. Reactivity is arguably an active ingredient in any SM intervention, but the involvement of the external observer and the novelty of being involved in a study introduces an additional threat to the internal validity of the research known as the Hawthorne effect. The Hawthorne effect suggests that the behavior recorded during a research study may not be representative of the participants’ day-to-day life (Ledford & Gast, 2018), but an artifact of being observed, the novelty of study participation, or researcher expectations (McCambridge et al., 2014). Furthermore, participants’ knowledge and/or beliefs about the purpose of and expectations for a study are the strongest variables associated with producing the Hawthorne effect (Adair, 1984). This is most likely to occur with adults and in short-duration studies (McCambridge et al., 2014). The significance of the Hawthorne effect is that there is an unknown portion of behavior change that may not be attributable to the intervention and leaves room for error in understanding the mechanism of action (Marrelli, 2007).
The idea that the Hawthorne effect may confound findings in studies of SM is not new, but has not been explicitly controlled for in any single-case studies (many of which have short data series). The application of obtrusive assessment procedures, reactivity, and a known purpose of the study are the reported limitations referred to within the SM literature (Briere et al., 2015; Mouzakitis et al., 2015; Seligson-Petscher & Bailey, 2006). The present study adds to the SM literature by designing controls for a portion of reactivity associated with the Hawthorne effect by masking the primary participants from the true purpose of the study. Ethical and human participants research practices were fully considered, approved, and employed, which ensured awareness of study participation. Classroom observers (mentor teacher and researcher) were necessary to confirm the validity of the data. Given these conditions, this study did not attempt to control for the types of reactivity associated with observation, measurement, and the novelty of being involved in research. However, the current study design attempted to control for the most salient factor associated with producing the Hawthorne effect, which relates to participant beliefs about researcher expectations and a desire to change behavior to meet the perceived expectations. Masking was appropriate for the current study given the adult population and short duration. These procedures are consistent with recommendations that finding less obtrusive procedures can corroborate results attained by traditional observation methods and may serve as a comparison to quantify the portion of behavior change attributed to the Hawthorne effect (Marrelli, 2007).
This study examined the effects of a multicomponent SM intervention consisting of PF and GS on a classroom management practice with pre-service teacher interns. Research questions include the following:
Method
Participants
Participants included interns (primary participants) and mentor teachers (interventionists and data collectors). Mentors were recruited before interns. Their data are reported first.
Mentor teachers
All general education teachers with an assigned intern were eligible for participation (N = 13). Interested mentor teachers contacted the first author via email to set up one-on-one meetings to discuss the study procedures and those with continued interest were consented by the first author (n = 2 in third grade, n = 3 in fourth grade, n = 5 in fifth grade). This study utilized five fifth-grade teachers. All mentors were Caucasian females. Two teachers shared a mentor position by each working two and a half weeks (half) of the 5-week session. Mentor backgrounds include master’s degrees (n = 3) and bachelor’s degrees (n = 2). Teaching experience ranged from 3 to 31 years and all had prior experience mentoring pre-service teachers.
Interns
Pre-service interns were the primary participants. Mentors described the research opportunity and provided the initial informed consent to the intern within the first few days of the session. Interested interns contacted the first author via email to set up a one-on-one meeting. Initial eligibility included (a) intern teacher status, (b) teaching in the summer school session, and (c) assigned to a consented mentor. Eligible and willing participants totaled four. After consent, a brief observation confirmed low use levels of opportunities to respond (OTRs) and behavior-specific praise. All four interns were Caucasian females between the ages of 22 and 24 years. Intern backgrounds were homogeneous. Each held an undergraduate degree in a field other than education (philosophy and religion, Intern 1; English, Intern 2; psychology, Interns 3 and 4) and were currently pursuing a master’s of art in education degree with an emphasis in elementary education. Each had a minimum 100 hr of field experiences in elementary schools. Intern 1 had previous field experiences in a second grade and gifted classroom. Intern 2 had spent time in the first- and fifth-grade classrooms along with previous experience as a summer school daycare teacher. Intern 3 experienced first- and second-grade field placements and had prior involvement at a preschool. Intern 4 had field experiences in first-, second-, and fourth-grade classrooms.
Masking primary participants
The protocol included masking to strengthen internal validity by controlling for a variable associated with the Hawthorne effect (Adair, 1984). All study procedures, including deception, were approved by the institutional review board (IRB) at the first author’s university. Participant orientation expressly stated that the purpose of the study was to investigate perceived satisfaction outcomes for mentoring meetings and relationships between the intern and the supervising mentor. These meetings were described as regularly scheduled, structured, and goal oriented. The IRB-approved informed consent document used deception and justified the faux study by indicating that mentors hold an important role in the education of a future teacher and the quality of the relationship may affect the development of a pre-service teacher. Interns believed satisfaction with the mentor–mentee relationship would be assessed through interviews and questionnaires at the end of the study.
The true purpose of the study and the intended data set of interest were initially hidden from the intern. Participants consented to study procedures initially but the rationale, purpose, and use of measures were masked. For example, terms such as “data collection” were used instead of SM. The data collection process was described as a structured way to scaffold mentoring and was held constant across conditions of baseline and intervention. Because participants have the tendency to change their behavior to align with study goals (McCambridge et al., 2014), the purpose of the study was not revealed at any time. Intern masking to the purpose of the study attempts to isolate effects of the multicomponent treatment package from reactivity, rather than interns potentially reacting to perceived goals of changing a classroom management practice. Interns understood the mentoring process to be the focus of the investigation. It was hypothesized that interns would be less likely to change behavior related to the target behavior, because they believed the researcher’s aim was satisfaction with the mentor–mentee relationship.
Informed consent and debriefing for the actual study occurred at the conclusion of the study. The first author explained the true purpose of the study and why masking was beneficial for the design. Interns agreed to participate in all study procedures in the initial consent document, but a second consent document was provided during the debriefing session seeking approval to use the SM data. All four interns gave permission to use the SM data. Mentors knew the true purpose of the study for data collection purposes.
Setting
The school was in a rural midwestern town with a population of about 17,500. Historically, the local university and school district partner during the summer to provide field experiences for interns in the master’s level education program. The public elementary school (Grades 3–5) had a population of 594 students during the previous academic year. The school was characterized by a 54.9% free and reduced lunch population and an 86% Caucasian student population. The research site had implemented school-wide positive behavior supports for more than 15 years. The study occurred during the 5-week summer school session. Procedures took place in four general education fifth-grade classrooms. Class sizes ranged from 23 to 27 students with total summer enrollment of 300 students. All students were invited to attend the full day summer school session where breakfast, lunch, and transportation were provided at no expense to the family. Students with qualifying reading scores participated in a separate gifted program.
Teaching interns from the master’s program were assigned to a 5-week placement with week-long teaching topic rotations for content areas of reading, writing, math, and science/social studies. Mentors were content area experts and taught one subject throughout the day. Intern teachers physically rotated through classrooms, teaching in four different classrooms settings every day. Mentors and interns used a commercially produced, district-adopted summer school curriculum across each topic, which included all supplies necessary to perform the lesson. Instructional scope and sequence were thereby controlled across teachers, interns, and each content areas. The typical size classrooms maintained the appearance of the school year with decorations, desks, technology (smartboard), and materials left in place by the prior teacher.
Measures
Dependent variables
The classroom management practices of behavior-specific praise (BSP) and OTRs were selected as readily observable and measurable practices that were free operants. BSP was defined as verbal praise that acknowledges a specific positive social or academic behavior. (e.g., nonexample: “Great job!”; example: “Thank you for keeping your hands to yourself in the hallway!”). OTR for this study was defined as the number of times the teacher provides academic requests that require more than one student in the classroom to actively respond at a time. This response could be verbal, written, or in the form of a gesture (e.g., nonexample: asking a question and students raise their hand for the teacher to call on them one by one; example: choral response, use of whiteboards, response cards, thumbs up/down). Individual responses were excluded based on the request from the summer school director. Interns often struggled to actively engage larger portions of the classroom in prior years and the school preferred to promote group OTRs.
At the training prior to intervention, the first author asked interns to choose either OTR or BSP to serve as their individual target behavior. Choice between the two practices was used to increase participant buy-in. All four interns independently chose to target BSP throughout the intervention condition. The data collectors measured the variables by systematic direct observation, during a specific 15-min time period of teacher-directed instruction (Ledford & Gast, 2018). Mentors used data collection sheets to record the frequency of the behaviors with tallies for each instance. Interns used a handheld counter to monitor the frequency of the behavior during intervention. This method was chosen based on the work of Simonsen et al. (2013), which found that tallying and using handheld counters for SM was most effective, but teachers preferred the handheld counter.
Data collectors also gathered information on time on task (TOT) to examine the impact on student outcomes. TOT was measured by momentary time sampling using planned activity check (PLACHECK) to report group behavior (Ledford & Gast, 2018). The data collector scanned the room and recorded the number of students on task at the end of a 2-min interval. If the student was observed as engaged in the task assigned by the teacher, he or she would be counted as on task (e.g., nonexamples: out of assigned space, playing with items at desk, gazing around the room; examples: pencil in hand working on assignment, looking at the teacher/board during teacher-led instruction, discussing with peers during group work). This cycle repeated until the 15-min data collection period ended resulting in eight data points. Data are reported as the percentage of intervals in which 80% or more of the students are on task. PLACHECK requires both observers to visually scan the room at the same speed to produce reliable results. To control for scanning speed, the mentor and first author chose a random grouping of four to five students in one area of the classroom.
Interobserver agreement (IOA)
IOA occurred a minimum 20% of data sessions for each condition and participant (Kratochwill et al., 2010). The mentor collected primary data and the first author collected IOA data. Both mentor and the first author tallied the frequency of the target behaviors for the 15-min time period in 2-min intervals (one 1-min interval). IOA data are reported by the percentage of intervals in agreement.
The first author assessed IOA for the intern’s target behaviors (OTR and BSP) and student TOT an average of 26% across participants. Percentage of sessions observed ranged from 23% to 40% in baseline, 20% to 33% in intervention, and 20% in maintenance. Agreement was calculated by dividing the number of intervals that reported the same frequency by the total amount of intervals (eight intervals) available in the 15-min time period and then multiplying by 100 to report the agreement as a percentage. IOA exceeded acceptable levels on intern target behaviors (OTR and BSP; 80% standard; Kratochwill et al., 2010) for the overall study (97.5%), all conditions (baseline 100%, intervention 93.1%, maintenance 100%), across target behaviors (BSP = 96%, OTR = 98.5%), and participants (1, 95%; 2, 96%, 3, 99%; 4, 100%). IOA met standards for TOT for the overall study (88%), all conditions (baseline 90.6%, intervention 84.3%, maintenance 88%) and across participants (1, 85%; 2, 88%; 3, 88%; 4, 91%).
Experimental Design
A multiple baseline across participants single-case research (SCR) design (Kennedy, 2005) was used to examine the effectiveness of SM on four (masked to study) interns’ use of a classroom management practice in the classroom. The design provided four opportunities to demonstrate effect, one for each onset of treatment; consistent effects across three of the four legs of the design would indicate a functional relation (Horner et al., 2005; Kratochwill et al., 2010; Maggin et al., 2013). Concurrent baseline data with stability determined onset of treatment for a randomly selected next participant. Subsequent introduction of intervention occurred randomly for remaining participants after a minimum of two data. The pattern repeated until all interns had entered. Prompting was a planned procedural change if the intern was nonresponsive. This decision was always made in consultation with the mentor. Phase changes occurred Tuesday through Friday to eliminate setting and mentor differences as a confound.
Procedures
Baseline
Interns were responsible for all instruction in the morning session and the mentor offered instructional support. Mentor, summer school director, and university supervisors advised per usual (PF, both written and verbal). The research team did not limit feedback on any topic. Prior to baseline condition, mentors received a 25-min group training on the operational definitions of OTR and BSP and data collection procedures. Mentors collected data on OTR, BSP, and student TOT. Data were recorded during a 15-min time period of teacher-directed instruction. Time slot remained constant throughout the study with each mentor, subject, and condition. OTR and BSP were both monitored for decision-making purposes.
Intervention
The intervention included SM, PF from the mentor, and GS. In this study, SM was the process of the intern observing and recording the frequency of the target behavior. PF from the mentor included verbal or written communication provided to the intern regarding their performance on the target behavior (BSP) and progress toward previously selected goal. GS involved the intern setting a target frequency for each day, discussing that number with the mentor, and recording it on a simple bar graph. Prior to starting the intervention, mentor and intern received a 20-min training on study procedures. Precautions were taken to ensure interns remained masked to the purpose of the study. During the intervention condition, the intern monitored the frequency of the chosen target behavior (BSP) with a handheld counter for a 15-min time period. During this time, the mentor teacher recorded the rates of OTR, BSP, and TOT for students. After the 15-min session, data collection stopped and the intern finished the lesson. Mentor and intern met to debrief about the target behavior after completing the lesson each day. The debriefing session was regularly scheduled (daily), structured (specific steps), and goal oriented (frequency count of practice). During the brief meeting (10–15 min), mentor and intern compared frequency counts, discussed discrepancies, and agreed on a final count to report. The mentor provided PF on the intern’s use of target behavior. Following this, the pair discussed progress toward previous goal, graphed the data, and set a goal for the next day.
Maintenance
Interns entered maintenance after a minimum seven data or maximum 10 data points in intervention. Mentor continued data collection on OTR, BSP, and student TOT every day. Interns did not collect data on their behavior, engage in the debriefing session with PF, or set next-day goals. Mentor and university supervisors provided PF as during baseline.
Comparison behavior
The intern self-monitored the frequency of one practice consistently during the intervention, but mentors and researcher collected data on both OTR and BSP throughout all conditions. The frequency data of the practice not chosen for the SM intervention served as a comparison to further explore the strength of the functional relation. All four interns chose to SM BSP throughout the intervention condition, which meant that OTR served as the comparison behavior. It was the expectation that the SM intervention would increase the targeted behavior (BSP), and the comparison behavior (OTR) would remain relatively stable throughout conditions.
Procedural Fidelity
The mentor used a daily checklist to report protocol adherence. The first three steps were observable during the teaching period and included setting timer, collecting frequency data on handheld device, and recording final count. The following steps occurred during the debriefing session, which included reviewing data, PF, graphing data, and GS. Mentors reported 100% adherence to intervention steps. To preserve the integrity of the private meeting, the research team did not collect reliability on Steps 4 through 10. The first author collected IOA for the first three steps in at least 20% of sessions in each condition for all participants and mentors. Reliability between mentor and first author resulted in 100% agreement.
Social Validity
Social validity of goals, procedures, and outcomes were assessed as a more comprehensive and valid measure (Snodgrass et al., 2018; Wolf, 1978) than a brief checklist. Goal validity was assessed with a one-on-one interview of the summer school director prior to implementation. The target behavior was also self-selected by interns. Feasibility of the procedures and effectiveness of the intervention were assessed after the study using electronic anonymous questionnaires and individual semistructured interviews. The questionnaire consisted of seven questions for the mentor teacher and 11 questions for the intern that prompted the individual to respond on a 1 (strongly disagree) to 5 (strongly agree) scale. Researchers reviewed data from the interviews to identify themes.
Analysis
Quality analysis
The first and second authors self-checked the study procedures against quality standards for What Works Clearinghouse (WWC, 2017). A doctoral student previously trained in the WWC standards served as a third independent rater. The external rater scored the study as meets standards using the WWC’s (2017) pilot single-case design standards.
Visual analysis
The first and second authors examined the visual representation of the data to determine the presence or absence of a functional relation by focusing on mean-level change, trend, and variability within and between phases, and immediacy of change at the intercept gap (Kratochwill et al., 2010; Ledford & Gast, 2018). Per current practices and recommendations, a masked visual analysis was then completed by three independent raters with no knowledge of the study (Byun et al., 2017; Vannest et al., 2018). Raters evaluated the data visually after a brief training of definitions (i.e., trend, overlap) and scored the data using a coding sheet. External raters also estimated within- and between-condition level, trend, variability, immediacy, and consistency. Decisions across replications (data patterns, immediacy of change, overlap, and consistency) were summarized against the three-fourths rule (Maggin et al., 2013). This served as the summative analysis to determine the existence of a functional relation. Data demonstrating a functional relation were then assessed for magnitude of effect by comparing the amount and consistency of change across all conditions of each participant (Kennedy, 2005; Ledford & Gast, 2018).
Statistical analysis
Effect size (ES) calculations can provide an additional source of data and are recommended in conjunction with the visual analysis so long as they reflect the design logic and include reference to context (Hott et al., 2015; Vannest et al., 2018; Vannest & Ninci, 2015). Nonparametric ES indices identify the quantity of change between A and B phase contrasts, as well as an omnibus ES across replications within the design. Tau and Tau-U statistical analyses were selected as a recommended nonparametric ES (Zimmerman et al., 2018). It is a robust indicator of the magnitude of effect using the logic of overlap consistent with visual analysis (Wolery et al., 2010). Tau-U has 95% the power of parametric tests when data are ideal for those analysis, power exceeds 100% when data are skewed and nonnormal, which are common characteristics of SCR data (Parker et al., 2011). Tau is the ES without adjusting for trend. Tau-U is the ES with trend controlled. ESs were computed using the free, online calculator (Vannest et al., 2016).
Results
Mentors collected direct observation data across all phases for intern and student behaviors (e.g., BSP, OTR, TOT, and SM use). All four interns selected BSP to target and the data are presented in Figure 1.

Intern rates of target behavior.
Within Participants
Intern 1 displayed low and stable rates of the target behaviors and stayed in baseline condition for five sessions (BSP: M = 0.8, SD = 1.10; OTR: M = 0.2, SD = 0.45). Starting intervention, she exhibited a level increase in BSP with small amounts of variability and a slight increasing trend, whereas OTR remained at low and stable rates (BSP: M = 6.88, SD = 1.13; OTR: M = 0.38, SD = 0.74). This first intern’s data are consistent with a demonstration of effect in the design, should the change be replicated. Intern 1 progressed through to maintenance prior to the conclusion of the study. She displayed similar levels of BSP in maintenance when compared with the intervention phase, with a small increase in level at the start of the condition, moderate variability throughout, and an eventual decreasing trend (M = 7.6, SD = 3.29).
Intern 2 exhibited a decreasing baseline trend for BSP and OTR with moderate variability with BSP and elevated variability with OTR. She stayed in baseline for eight sessions (BSP: M = 1.38, SD = 1.51; OTR: M = 3.75, SD = 4.09). BSP increased level immediately in intervention and maintained high levels with moderate variability and one overlapping data point (M = 9, SD = 4.6). Although Intern 2 displayed a sudden decrease on the second day of intervention, the mentor advised against implementing prompting by explaining “unusual circumstances” of that data collection period. Data returned to high levels the following day. These data demonstrated a second effect, or consistent replication. OTR did not display a change in level with moderate variability with a decreasing trend (M = 2.5, SD = 3.25).
Intern 4 showed relatively low rates of BSP and OTR with moderate variability and stayed in baseline for 11 sessions (BSP: M = 0.64, SD = 0.67; OTR: M = 1.69, SD = 2.25). Intern 4 was the third participant to enter the intervention and demonstrated an immediate change in BSP for two sessions before dropping to a frequency of zero. The mentor for participant 4 reported the intern as “drifting” from a number of instructional expectations including those targeted in the study and expressed a desire for a naturalistic adaptation of “prompting” or “pre-correcting” to help the intern focus. Prior to the teaching period, the mentor reminded the intern of the goal and to engage in BSP. Data reflecting this prompting condition is distinguished as an additional phase in the design. After the introduction of prompting, there was an increase in BSP level during the remaining sessions (M = 3.83, SD = 1.94). OTR did not have an immediate level change at the start of intervention, but slightly increased with the addition of prompting (M = 0.66, SD = 1.21).
Intern 3 demonstrated low and stable levels of BSP with slightly higher and more variable rates of OTR in a decreasing trend. Baseline lasted 13 sessions (BSP: M = 0.23, SD = 0.6; OTR: M = 0.82, SD = 1.17). In intervention, she had an immediate change in BSP level, maintained moderate levels with little variability, and displayed no overlap between conditions (M = 6, SD = 1). In comparison, OTR experienced a slight jump at the start of intervention, but displayed a decreasing trend and relatively low levels. Intern 3’s data demonstrate a third effect.
Across Participants
There were no statistically significant differences across any participant’s baseline data (e.g., all demonstrated equally low rates of baseline data). The mean rate of BSP across participants was <1 per 15-min interval (0.76). These scores meet an operational definition of low rates of BSP compared with recommended rate of six praise statements per 15 min (Myers et al., 2011). Rates of OTR were consistently low across participants during baseline (M = 1.62), in comparison with the suggestion of at least three per minute (Stichter et al., 2009). Trend in baseline for BSP was assessed visually and statistically using a pencil test with an extended celeration line and Tau-U. Increasing trend exceeding a threshold for correction of above .10 (−.4, .78, .01, and .16) occurred in three of the four baselines of BSP. At the onset of the SM intervention, each of the four participant interns demonstrated an immediate level increase in BSP. These changes were statistically significant at the .01 level. The nontarget behavior (OTR) remained similar to baseline levels.
Tau-U
Overall, for four participants in the design, 235 pairwise comparisons contributed data, producing an omnibus Tau-U value of 1 where data are weighted by the inverse of the variance. This is interpreted as 100% of treatment sessions demonstrated improved performance, which are attributed to the intervention. The confidence intervals (CIs) for these data reflect variability and phase length, producing CI = [0.83, 1]. This is interpretable as a narrow band, providing evidence of consistency and confidence that the effect is large and not a false positive. Individual legs in the design for each of the four participants produced Tau-U values ranging between .88 and 1. When ties were counted, the Tau-b values ranged between .91 and 1. These statistical analyses agree with the masked visual analysis of consistent treatment effects and large change in behaviors for individual and the overall designs. All raw data are available from the first author.
TOT
Momentary time sampling of TOT for a random sample of five children in each class produced variable baseline levels (percentage of intervals with 80% students on task, M = 48%, 31%, 52%, 48%). TOT during intervention reflects 88%, 91%, 87%, 76% of intervals with 80% or more of students on task. Student outcome data produced Tau-U values ranging between .26 and .71. Mean-level changes showed improvements between 28% and 60%. However, the first and second authors do not believe a functional relation was established between teacher use of BSP and on-task behavior based on the variability, immediacy, and trend in the data. As illustration, baseline TOT ranged by participant between 16% and 87%, 0% and 63%, 0% and 87%, and 0% and 100%. Data are displayed in Figure 2.

Student TOT.
Social Validity
The research team addressed the three components of social validity as described by Wolf (1978): goals, procedures, and outcomes (Snodgrass et al., 2018).
Goals
Based on prior experience with novice teachers, the summer school director confirmed the validity of the goals of the study. She expressed that the practices chosen aligned with building priorities as a school implementing positive behavioral interventions and supports (PBIS). When given the option to choose OTR or BSP to monitor, all interns said they needed to increase BSP and expressed interest in doing so.
Procedures
Mentors and interns found the procedures of the SM intervention feasible and appropriate. Interns and mentors either somewhat agreed or strongly agreed the intervention was easy to implement (M = 4.5, M = 5). Interns somewhat disagreed that it was difficult to self-monitor while teaching (M = 2). When interviewed, all interns and mentors agreed that the intervention was reasonable and practical for practicing and pre-service teachers. Three of the interns specifically mentioned the ease of collecting data with a handheld counter. However, one intern felt uncomfortable SM while her mentor observed.
Outcomes
Overall, the intervention was considered to be effective. Interns strongly agreed that the intervention increased their awareness of their teaching practices (M = 5) and at least somewhat agreed that this increased awareness changed their teaching practices (M = 4.5). Correspondingly, interviews revealed that the intervention made the mentors more aware of their own teaching practices. Three of the four interns asked to keep the counter and expressed a desire to self-monitor in the future. All mentors strongly agreed that they would be willing to self-monitor or recommend it to another teacher (M = 5). Finally, mentors and interns at least somewhat agreed that the intervention benefited the students (M = 4.67, M = 4.5).
Discussion
The current study examined the effects of a multicomponent SM intervention on the rate of behavior-specific praise with pre-service teacher interns and students in masked conditions. After relatively low and stable rates of BSP and OTR in baseline, all four interns increased their use of BSP with the onset of the SM intervention. Simultaneously, nontargeted OTR remained in the range of the baseline condition. Consistent effects across three participants demonstrated a functional relation, and data indicated agreement between visual and statistical analyses. One intern required prompting prior to SM condition, and with prompting demonstrated a functional relation between the intervention and an increase in BSP. The results of this study support the use of an SM intervention including PF and GS as a promising way to increase the use of BSP with pre-service teachers, and is consistent with previous literature in the field (Alexander et al., 2012; Hager, 2012; Keller et al., 2005). This study further contributes to the SM knowledge base by demonstrating effects with teacher interns in masked conditions. The selected components (SM, PF, and GS) and the collaborative nature of the present treatment package also relate to the teacher coaching literature, which has shown positive effects for behavior change with pre-service and in-service teachers (Coogle et al., 2015; Kraft et al., 2018; Stormont et al., 2015).
Experimental investigations require researchers to take great care in choosing a design that rules out confounding variables as a possibility for the behavior change (Ledford & Gast, 2018). Additional steps were taken in the current methodology to allow for a greater sense of confidence when attributing the results to the intervention. The interns were masked to the true purpose of the study and the primary data set of interest. Participants may have been less likely to change their behavior related to the SM component of the intervention, because they believed the purpose of the study and researcher expectations related to the mentor and mentee relationship. To provide additional control, an additional practice other than the target behavior was monitored during the intervention phase. Rates of BSP increased with the onset of intervention, whereas OTR rates remained relatively close to baseline levels. Previous work has suggested a possible positive correlation between praise and OTR (Sutherland et al., 2002), which may explain any slight changes in the rates of OTR during the intervention phase. The lack of effect with a potentially related practice further supports that the increase in BSP is due to the intervention. To the authors’ knowledge, no studies of teacher behavior within the SM literature include an unrelated comparison practice or masking in the protocol.
In addition to teacher outcomes, student TOT was collected as a distal measure of the effect of an increase in the use of BSP. Research indicates that teacher-delivered BSP shows 73% to 100% of treatment sessions with improvement on student performance outcomes using Tau-U ES (Royer et al., 2019). The data obtained in the present practitioner-focused study are inconclusive related to student outcomes with ESs showing 26% to 71% improvement. In a review on teacher SM with behavioral practices, Rispoli et al. (2017) noted that only 47% studies reported student outcomes with both positive (Allen & Blackston, 2003; Plavnick et al., 2010; Szykula & Hector, 1978; Workman et al., 1982) and mixed results (Bingham et al., 2007; MacSuga & Simonsen, 2011; Mouzakitis et al., 2015; Reinke et al., 2008). The mixed results in the current study may demonstrate an actual lack of effect, but also may be the result of insufficient sensitivity in the measurement tool or suggest a certain degree of latency between teacher change and student behavior change. Unlike the large group setting (23–27 students) in this investigation, the studies that reported student outcome data in the literature review ranged in total group size from one to eight students, with those reporting positive effects ranging from two to five students (Rispoli et al., 2017). It is possible that rates of BSP did not rise to the level necessary to show an effect with this larger group. Although, mean BSP levels of three interns met the recommendation of at least six per 15-min session (Myers et al., 2011).
Social validity data suggest the SM intervention was viewed as both feasible and effective by the interns and mentors. The interns’ affinity for the handheld counter was consistent with prior research (Simonsen et al., 2013), and the desire to keep the counters supports the usability and perceived efficacy of the intervention. Survey and interview data suggested an increased level of awareness in the interns, but it also revealed that collecting data and providing feedback made mentors more aware of their own teaching practices. More information is needed about the potential residual effects for those collecting data and proving the PF.
Limitations
The present results should be interpreted with an understanding of the limitations of the study. To begin with, all mentors had prior experience working with future teachers. This may have benefited them in their interactions with the interns. Second, allowing interns to choose between OTR and BSP prior to the intervention condition may have influenced the rate of the comparison behavior. Intern 2 chose BSP to monitor, but expressed an interest in increasing OTR after seeing her baseline scores. Due to the masked intervention condition, mentors did not actively limit her use of the practice. Increased awareness of the comparison practice may have been responsible for any variability in OTR. Finally, generalization probes in alternative settings were not conducted. It is possible that the increase in BSP did not generalize across settings.
Implications and Areas for Future Research
Data suggest the SM intervention including PF and GS, may be helpful in improving practices with pre-service teachers during field experience placements. Future studies should investigate the possible residual effects for mentors who utilize this method with a pre-service teacher. Moving forward, researchers should consider additional ways to control for threats to internal validity, such as implementing masked protocols or collecting data on comparison behaviors if available. Efforts made to strengthen experimental control will increase the rigor of SCR methodology and further investigate the level of behavior change that can be attributed to the intervention. In addition, the inconclusive results relating to the student TOT data after an increase in intern use of BSP suggest that additional research should include student outcomes. Inconsistent reporting of student data and mixed results in current literature on teacher behavior speak to the difficulty in selecting and measuring the appropriate student outcome. Researchers should focus on distal outcome measures in teacher interventions by exploring student growth. Finally, additional research should consider social validity more completely. A recent review of the SCR literature indicated that 26.8% of the articles reported the results of a social validity assessment and only 6.5% of the total studies addressed the three constructs: goals, procedures, and outcomes (Snodgrass et al., 2018). To meaningfully address the research to practice gap, researchers must ensure practices are viewed as socially valid.
Conclusion
This study identified a functional relation between the treatment package and increased use of an evidence-based practice (EBP) for three of the four masked participants as identified by masked visual analysts and consistency between visual analysis and statistical analysis. There were no contraindicated effects. The fourth participant required additional prompting to engage in the treatment package. Social validity across all three constructs was high. Use of a masked visual analysis and a random assignment during a masked-condition study is infrequently seen in the current literature, but may provide a template to encourage scientific exploration in this area, particularly when outside factors may influence treatment effects.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
