Abstract
The practice of screening students to identify behavioral and emotional risk is gaining momentum, with limited guidance regarding the frequency with which screenings should occur. Screening frequency decisions are influenced by the stability of the constructs assessed and changes in risk status over time. This study investigated the 4-year longitudinal stability of behavioral and emotional risk screening scores among a sample of youth to examine change in risk status over time. Youth (N = 156) completed a self-report screening measure, the Behavioral and Emotional Screening System, at 1-year intervals in the 8th through 11th grades. Categorical and dimensional stability coefficients, as well as transitions across risk status categories, were analyzed. A latent profile analysis was conducted to determine if there were salient and consistent patterns of screening scores over time. Stability coefficients were moderate to large, with stronger coefficients across shorter time intervals. Latent profile analysis pointed to a three-class solution in which classes were generally consistent with risk categories and stable across time. Results showed that the vast majority of students continued to be classified within the same risk category across time points. Implications for practice and future research needs are discussed.
Screening is recommended for use in schools as an initial step for providing early preventive and intervention services. In particular, screening for emotional and behavioral risk has been advocated for based on the high point prevalence rates for mental health problems among youth (Essex et al., 2009) and because of the known benefits of early identification and intervention (Glover & Albers, 2007). Screening for behavioral and emotional risk is differentiated from screening from mental health disorders in that the interest is on screening for early symptoms of disorder as opposed to identifying students who meet diagnostic or special education classification criteria (Kamphaus, 2012). Significant progress has been made in regard to academic screening efforts within schools, with an increasing number of schools using academic screening and progress monitoring tools to determine who may benefit from additional educational services. However, the process of screening for behavioral and emotional risk has lagged behind academic screening and there are still uncertainties that need to be addressed through future research (Cook, Volpe, & Livanis, 2010). The purpose of this longitudinal study is to investigate the stability of behavioral and emotional risk screening scores among a sample of youth to examine change in risk status over time.
Frequency of Behavioral and Emotional Screening
One of the major uncertainties regarding behavioral and emotional screening implementation is when and how often to screen (Dowdy, Furlong, Eklund, Saeki, & Ritchey, 2010). The call for systematic and continuous screening in schools has been made in an effort to preempt the “wait to fail” model typical of the special education process (Caldarella, Young, Richardson, Young, & Young, 2008; Husky, Sheridan, McGuire, & Olfson, 2011); however, few practice guidelines with regard to frequency have been offered, and at times, the suggestions are contradictory. Some scholars have suggested that behavioral screening data need to be collected on a “proactive and continuous basis” (Chafouleas, Kilgus, & Wallach, 2010, p. 250), such as screening every child multiple times throughout the year (Ennis, Lane, & Oakes, 2012). Others have focused on critical developmental and institutional time points, such as screening during the transition from elementary to middle school, when students may be at heightened risk for emotional distress, but when early symptoms may still be addressed prior to reaching a diagnostic threshold (Stoep et al., 2005). Suggestions to screen all students at the beginning of each school year have been offered (Spielberger, Haywood, Schuerman, & Richman, 2004), as well as suggestions to screen for emotional and behavioral risk when academic concerns present (Dowdy et al., 2010). Severson, Walker, Hope-Doolittle, Kratochwill, and Gresham (2007) recommend that multiple-gating screening assessments, such as the Systematic Screening for Behavior Disorders (Walker & Severson, 1992), be conducted twice a year, once at the beginning of the school year when teachers become familiar with the student’s behavioral functioning and once at the beginning of the second semester to capture potential changes in the student’s behavior. Unfortunately, the overall recommendations in the field are inconsistent and, at times, appear to be made based more on practicality than on empirical data.
One feasible option is to conduct screenings on an annual basis, such as recommended by the U.S. Preventive Services Task Force for the screening of depression by primary care physicians (Sharp & Lipsky, 2002). Annual screenings align with the current school context, where each year students transition to a new grade with different personnel responsible for their growth. Annual screenings, early in the school year, may provide the opportunity for early identification of emotional or behavioral problems that have the potential to affect their academics that year. Decisions based on screening data can then be used to direct early intervention or prevention services for that academic year. However, the merit of annual screening has not yet been subjected to systematic evaluation.
Screening for behavioral and emotional risk needs to be conducted more often when there is substantial change in risk across time (i.e., high rates of transiency, compounded risk in adolescence as compared with elementary age; Caldarella et al., 2008; Walker, 2010). However, particularly in a time of limited resource allocation, if a large majority of students continue to be classified into the same risk category on each subsequent rescreening, then it may not be cost-effective to rescreen as frequently (Nease & Malouin, 2003). On the other hand, if the negative outcomes associated with not treating the few cases that change into a more elevated risk status is high, then rescreening may be needed. Examining the stability of screening scores is a necessary early step in determining an optimal screening schedule.
Stability of Behavioral and Emotional Screening Scores
The stability of screening scores is important to consider, as educational decisions are often made based on the results. Screening results often lead to determinations regarding who should receive additional follow-up preventive or early intervention services. Elevated scores may be the catalyst for further assessment and intervention, whereas scores within the normal range may lead to the denial of additional services or even the dismissal of services for a student previously in the elevated range. Additionally, the rationale for screening for behavioral and emotional risk is grounded on the assumption that behavioral and emotional problems are at least moderately stable across time and that there are potential negative long-term effects if the problems are left untreated (Essex et al., 2009).
When determining how often to conduct systematic screenings, the stability of scores merits consideration. Decades of previous research have provided compelling evidence for the relative stability of behavioral and emotional problems across time. Reviews of the literature (Loeber, 1982; Olson, Schilling, & Bates, 1999; Stemmler & Losel, 2010) indicate that when high levels of externalizing problem behaviors are established in childhood, they often persist over time and are unlikely to decrease. Providing further evidence for the stability of externalizing problems, one 19-year longitudinal study demonstrated that a significant number of 4- to 6-year-old boys identified with an externalizing behavior profile remained so at 23 years of age (Asendorpf, Denissen, & van Aken, 2008). In general, externalizing problems have been found to be more stable than internalizing problems (Essex et al., 2009). However, most research indicates that a child with a comorbid profile (endorsing both internalizing and externalizing problems) is the most stable over time, presents the greatest functional impairments, and is in need of the most services (Essex et al., 2009). Lane, Oakes, and Menzies (2010) suggested that the population of students labeled with severe emotional and behavioral problems make limited academic progress, again highlighting the importance of early identification and treatment of risk factors prior to potentially stable maladaptive trajectories.
The short-term stability of scores has been investigated for a number of popular social-emotional screeners. In a review of screening instrumentation, Levitt, Saka, Romanelli, and Hoagwood (2007) reported stability coefficients for a number of behavioral and emotional screening instruments. For example, the 4- to 6-week stability for the Pediatric Symptom Checklist (Jellinek, Murphy, & Burns, 1986) was .80 (Navon, Nelson, Pagano, & Murphy, 2001); the 4- to 6-month informant stability coefficient for the Strengths and Difficulties Questionnaire (Goodman, 1997) was .62 (Goodman, 2001; Goodman, Meltzer, & Bailey, 1998); and the 1-month stability coefficients for the Systematic Screening for Behavior Disorders (Walker & Severson, 1992) ranged from .74 to .88. Other commonly utilized screening tools, including the Behavioral and Emotional Screening System (Kamphaus & Reynolds, 2007), reported short-term (median test interval 20 days) stability coefficients in the range of .80 to .91. The test–retest reliability has been extensively studied for the Student Risk Screening Scale (Drummond, Eddy, Reid, & Bank, 1994) with stability coefficients ranging from .60 to .86, and results providing evidence of higher stability coefficients in the short term (13-19 weeks) and moderate stability for longer time intervals (e.g., 22-84 weeks).
Current Study
The long-term stability of screening scores has not yet been subjected to investigation but can help answer how often screenings should be conducted. The current study was designed to test the stability of self-report screening scores over a 4-year period for a widely used self-report screener, the Behavioral and Emotional Screening System Self-Report Child/Adolescent Form (BESS Student; Kamphaus & Reynolds, 2007). Annual screenings were conducted to provide an initial investigation into the stability of screening scores. The results of this study could provide an initial starting point for more detailed investigations. Both continuous and categorical methods of analysis were used to offer complementary information reflecting both the dimensional and categorical approach to screening for emotional and behavioral risk in youth (Bilancia & Rescorla, 2010). More specifically, we sought to address the following research questions:
How stable are continuous scores of overall risk across a 4-year period?
How stable are categorical classifications of overall risk (i.e., Normal risk, Elevated risk, and Extremely Elevated risk) across a 4-year period?
How stable are continuous scores of subtypes of risk (e.g., internalizing and externalizing) across a 4-year period?
Do general patterns in the consistency of screening profiles over time emerge across a 4-year period?
Method
Participants
All 8th-grade students at one school were invited to participate, regardless of their general education or special education status. Following active parental permission and human subjects approval, a sample of 156 students (56% female) from one junior high school in central California participated in this longitudinal study, which represents a 52% participation rate for that grade from the school. Data were collected in four measurement waves at 1-year intervals in the fall of each academic year beginning in 2009. Self-report surveys were completed when the youth were in 8th grade (Time 1), 9th grade (Time 2), 10th grade (Time 3), and 11th grade (Time 4).
At Time 1, the average age of students was 13.27 years (SD = 0.54). Of the 156 students, 46% reported that they were Latino/a, 44% White, 9% African American, and the remainder reported they were of Asian, Pacific Islander, or indigenous American backgrounds. This is similar to demographic information for the school from which the sample was originally invited to participate, which had 41% Latino, 50% White, 2% African American, 5% Asian or Pacific Islander, and the remaining other or multiracial. The school had 36.7% of students receiving free or reduced lunch, which is lower than the average for the county or state.
Of the 156 children who were screened at Time 1, 72% of the original sample (N = 113; 54% females) remained across all four time points. At Time 2, 124 students participated. A total of 32 students were lost at the Time 2 follow-up due to moving out of the district to attend a private school or having moved (n = 25) or declining to participate (n = 7). At Time 3, 123 students participated. Some of these (n = 7) were students who were students who returned to the district at Time 3. Others (n = 8) who had completed the survey at Time 1 and Time 2 were not available at Time 3 because they were now in independent study, yielding a net loss of one student from Time 2 to Time 3. At Time 4, an additional 10 students were lost. Students who dropped out of the study had significantly higher T scores in 8th grade (M = 57.52, SD = 12.18) than those who remained for all 4 years (M = 50.70, SD = 10.51). An attrition analysis examining for differences among the students who dropped out of the study versus those who remained across all four time points revealed significant differences, t(153) = −3.52, p < .05, on behavioral and emotional risk (BESS Student scores).
Procedure
Data were collected as part of a longitudinal study associated with the evaluation of a Safe Schools Healthy Students (SSHS) project (Sharkey et al., 2011). As part of the SSHS project at the junior high school and high schools, there was increased access to universal prevention programming related to alcohol and drug use, as well as increased availability of mental health services. However, students were not referred for services based on their BESS Student scores, as the mental health program was specifically for students who experienced a traumatic event. Because of confidentiality issues, it is not known if any of the participants in the longitudinal study accessed mental health therapy services.
In the fall of each academic year for four consecutive years, university personnel administered the self-report surveys to participating students in a group or individual format during the regular school day. Students were excused from class and informed that participation was voluntary and that their responses would be anonymous after linking their data based on a unique identification number. Students placed their names on a cover sheet that was removed from the survey after linking their data using the unique identification numbers. After Time 1, with the assistance of school district personnel, students were tracked into their respective high schools. Students who subsequently enrolled in schools outside of the district (e.g., private school) were not available for inclusion in the following years, despite the attempts of the researchers to invite the school to participate. The students enrolled at a private school at Time 2 represented the largest group of students (n = 24) who dropped out of the study. In addition, as students moved through high school, some students were lost to follow-up after repeated attempts because they enrolled in nontraditional education programs, such as independent study (n = 7), an alternative/continuation high school (n = 4), or home-hospital (n = 2), to complete their high school requirements. This became more common in the 10th (Time 3) and 11th grades (Time 4). The remaining missing students were lost to follow-up despite repeated attempts to locate them each year.
Instrument
BESS Student
The Behavior Assessment System for Children-2 Behavioral and Emotional Screening System Student self-report form (BESS Student) is a 30-item self-report behavior rating scale designed to measure behavioral and emotional risk among students in 3rd through 12th grade (Kamphaus & Reynolds, 2007). Although the BESS Student was made available for students in both English and Spanish, only the English version was selected by the students in this study.
Students reported on their behavioral and emotional functioning using a four-point response scale (never, sometimes, often, almost always). The sum of the item raw scores is transformed to a total T score (mean = 50, SD = 10), in which higher scores reflect more problems. Based on their T scores, students were classified as having a Normal (20-60), Elevated (61-70), or an Extremely Elevated (71 and higher) level of risk. Exploratory and confirmatory factor analytic work suggests that the BESS Student measures four main constructs: inattention/hyperactivity (e.g., difficulty sitting still), internalizing problems (e.g., depression), school problems (e.g., attitude to school and teachers), and personal adjustment (e.g., social stress; Dowdy et al., 2011).
The psychometric properties of the BESS Student version are generally acceptable, having good split-half reliability (.90-.93) and test–retest reliability (.80) and moderate correlations with other measures of behavioral and emotional problems including the Achenbach System of Empirically Based Assessment Youth Self Report (.81), Conner’s Rating Scales (.65), the Revised Children’s Manifest Anxiety Scale (.53), and the Children’s Depression Inventory (.48; Kamphaus & Reynolds, 2007). For this sample, the BESS Student had acceptable internal consistency with Cronbach’s alpha ranging from .76 (Time 4) to .82 (Time 1).
Statistical Analyses
Analyses were conducted using SPSS Statistics Version 18.0, Mplus Version 7.0 (Muthén & Muthén, 1998-2011), and R Version 3.0.0 (R Development Core Team, 2013). Stability is generally estimated by administering the same measure to a group of respondents over a period of time. The test and retest scores are then correlated to produce a stability coefficient. Stability coefficients were analyzed for both dimensional, standardized T scores (M = 50; SD = 10), and categorical scales (Normal, Elevated, Extremely Elevated) produced from the BESS Student. Bivariate Pearson product–moment correlations were used to assess the relations between the dimensional scores at each of the four consecutive time points, whereas Spearman Rho correlations were used to assess the categorical correlations. Additionally, to investigate whether certain types of problems may be more stable than others, researchers calculated subscales based on factor analytic work (Dowdy et al., 2011). Specifically, the mean of the items that loaded onto each factor was calculated to form four subscales: externalizing problems (inattention/hyperactivity), internalizing problems, school problems, and personal adjustment. Steiger’s Z-tests (Steiger, 1980) for comparing dependent correlations were conducted in R version 3.0.0 (R Development Core Team, 2013) and were used to determine if certain subscales were more or less stable across time, and to compare stability coefficients across longer and shorter time intervals. Transitions between Normal, Elevated, and Extremely Elevated categories from Time 1 to subsequent time points were examined. Interpretation of correlation coefficient effect size was guided by Cohen (1988): r≤ .10 is small, r = .30 is moderate, and r≥ .50 is large.
Data were screened for valid response patterns and cases were removed when invalid response patterns were detected (i.e., answering never for all items). Pairwise deletion was used for all bivariate correlation analyses. Thus, the sample sizes varied for the bivariate correlations depending on the variables and time points that were included in the analysis.
A latent profile analysis (LPA) was conducted to determine key patterns that emerged among the continuous BESS Student items across four time points. An advantage of LPA with longitudinal data is that it can identify common underlying patterns that emerge in the data without imposing any fixed structure (Lanza & Collins, 2006; Nishina, Bellmore, Witkow, & Nylund-Gibson, 2010). LPA provides a general description of the most salient change patterns that exist in the population, providing another way to study the stability in scores over time.
All LPA models were specified in Mplus Version 7.0 (Muthén & Muthén, 1998-2011). Full information maximum likelihood estimation was used, which allowed for all participants to be included that had data on at least one time point (Enders, 2010). To control for attrition that may have occurred by participants who had higher risk for behavioral or emotional problems, we included a covariate that serves as a proxy for this potential missing data mechanism. The covariate of self-reported experiences with fights at school at Time 1 was included, and the LPA results presented controlled for this variable. LPA models with differing number of classes were fit, ranging from 1 to 5 classes. The final LPA model was determined by using a combination of model fit indices and the overall substantive fit of the latent classes. These fit indices include the Akaike information criterion (Akaike, 1987), Bayesian information criterion (BIC; Schwartz, 1978), and adjusted BIC (Sclove, 1987), where the model that yields the smallest values among these indices is considered the best fitting model. Moreover, the adjusted Lo–Mendell–Rubin likelihood ratio test (adjusted LMR-LRT; Lo, Mendell, & Rubin, 2001) and the nonparametric bootstrap likelihood ratio test (BLRT) are commonly used to test nested models (Nylund, Asparouhov, & Muthén, 2007). Specifically, they compare the K− 1 class model with the K class model. Significant p values suggest that the K class model fits the data significantly better than the K− 1 class model.
Results
Stability Coefficients
Stability coefficients were computed across all 4 years to compare each screening time point with each subsequent screening time point. All correlations were statistically significant (p < .001) and are presented in Table 1. Long-term stability of dimensional T scores ranged from .46 (Time 1 to Time 4) to .71 (Time 1 to Time 2), and stability coefficients for categories ranged from .40 (Time 1 to Time 3) to .51 (Time 2 to Time 3). When comparing the dimensional stability correlations in Table 1, results showed dimensional stability coefficients were significantly larger across shorter time intervals, ps < .05 (e.g., r = .71 compared with r = .46, p < .001).
Stability Coefficients for Dimensional, Categorical, and Factor Scores Over Time.
Note. Results were reported for all students who had data for the two time points being compared. All correlations significant at p < .001.
Stability coefficients for school problems ranged from .56 (Time 1 to Time 2) to .26 (Time 1 to Time 3), and stability coefficients for personal adjustment ranged from .63 (Time 1 to Time 2) to .43 (Time 2 to Time 4). Externalizing and internalizing problems were consistently not significantly different from each other with regard to stability across time, ps > .05, 1 with externalizing stability coefficients ranging from .62 (Time 1 to Time 2) to .43 (Time 1 to Time 3) and internalizing stability coefficients ranging from .63 (Time 2 to Time 3) to .33 (Time 1 to Time 4).
Change in Risk Status
When examining students’ classification into risk status categories, patterns of transition across time remained relatively stable (see Table 2). Of the students who were classified as Normal at Time 1 (n = 97), 92.8% (n = 90) remained classified as Normal at Time 2, 88.0% (n = 88) remained classified as Normal at Time 3, and 89.1% (n = 82) remained classified as Normal at Time 4. When movement did occur, it tended to be into the Elevated category (6.2% at Time 2; 9.0% at Time 3; 9.8% at Time 4). There were few students who transitioned from a Normal classification into the Extremely Elevated category, with only one student who was initially classified as Normal being classified as Extremely Elevated in Time 4. Overall, this indicates that the large majority of students initially classified as Normal remained so across the four points of inquiry.
Transition Matrix With Number and Percentage of Change Across Categories.
Note. E = elevated; EE = extremely elevated.
Of the students classified as Elevated at Time 1 (n = 15), 66.7% (n = 10) transitioned to the Normal category at Time 2, whereas 26.7% (n = 4) remained stable in the Elevated category. One student transitioned from Elevated to Extremely Elevated at Time 2. At Time 3, 61.5% (n = 8) of the initially classified Elevated students were in the Normal category, 30.8% (n = 4) remained in the Elevated classification, and the remaining 7.7% (n = 1) were classified as Extremely Elevated. By Time 4, 66.7% (n = 8) of initially labeled Elevated students were in the Normal category, and 33.3% (n = 4) were still in the Elevated category, and no students transitioned to the Extremely Elevated risk status. This demonstrates that a majority of students initially classified as Elevated either remained so across time, or, more commonly, they transitioned into the Normal classification.
Of the students classified as Extremely Elevated at Time 1 (n = 9), 33.3% (n = 3) transitioned to the Normal category at Time 2, whereas 44.4% (n = 4) transitioned to the Elevated category; 22.2% (n = 2) remained in the Extremely Elevated category. Nearly the same pattern held true across the remaining two assessment cycles. Overall, a minority of students initially classified as Extremely Elevated remained so across all four time points.
Latent Profile Analysis
Model fit indices for the two through five class models are presented in Table 3. The BIC as well as substantive interpretation pointed to the three-class solution (see Figure 1). Previous research has indicated that the BIC is the best and most consistent indicator of latent classes (Nylund et al., 2007). Although the LMR-LRT p value suggested a two-class model, the three-class model, which the BIC supported, had higher entropy and provided a more nuanced story of the trajectories and was the one we decided was best. The ordered three-class solution in Figure 1 shows the stability of the BESS Student T scores over time. Class 1 included 46.7% of the sample, Class 2 included 45.3% of the sample, and Class 3 included 8.0% of the sample. Students who were classified into Latent Class 1 had mean BESS Student T scores that hovered around 45 across the four longitudinal time points. Specifically, this class is made up of students who remain in the Normal range over time. Class 2 also included students who remained in the Normal range over time, although they are more elevated in their risk relative to Class 1 with mean T scores that hovered around 55. Last, Class 3 included students with the most Elevated and Extremely Elevated behavioral and emotional risk, with mean T scores that consistently hovered around 70 over time.
Fit Statistics for the LPA With One to Five Latent Classes.
Note. BLRT = bootstrap likelihood ratio test; LL= log likelihood; LMR LRT = Lo-Mendell-Rubin likelihood ratio test; LPA = latent profile analysis. Bolded values indicated the model that the fit index indicated was best.

BESS Student T Score mean plot for the three-class LPA solution.
Discussion
As behavior and emotions may vary across time and across contexts (Mash & Dozois, 1996), it is important to recognize that assessing behavioral and emotional risk at any one point in time may miss some students due to the transient nature of their symptoms. These results, however, suggest that self-reported screening scores for behavioral and emotional risk are at least moderately stable across time. Stability coefficients were similar to previous research investigating the test–retest reliability of screening scores (e.g., Ennis et al., 2012), although the time intervals for the current study were considerably longer. Findings were also consistent with prior research indicating that behavioral and emotional problems are relatively stable over time (e.g., Stemmler & Losel, 2010). However, findings that externalizing and internalizing problems were similar with regard to stability were unexpected as prior research indicates that externalizing problems are generally more stable (Essex et al., 2009). The small number of items assessing each construct, the small sample sizes, the reliance on self-report data, as well as the limited psychometric information in support of subscale factor scores may have contributed to these findings comparing the stability of externalizing and internalizing problems.
When examining movement across categories, findings clearly indicate that the majority of students have self-reported Normal levels of behavioral and emotional risk, and remain so across a 4-year time period. This aligns with expectations that the majority of students are not in need of specialized services and fits well within current multitiered intervention frameworks designed to offer individualized assistance to a small minority of students at the highest level of risk (Severson et al., 2007). When movement did occur, students tended to transition to a category most similar to their previous category (e.g., Normal to Elevated or Elevated to Extremely Elevated) as opposed to more drastic changes. Additionally, findings suggest that some students who were initially classified in the Normal range but then transitioned into the Elevated category transitioned back to the Normal category by the final assessment point. In fact, only one student classified in the Normal category at the initial screening was classified in the Extremely Elevated category at Time 4. There was more movement across categories for the students who were classified as Elevated or Extremely Elevated, within a given year. Fortunately, approximately two-thirds of these students tended to move to the Normal range. This could indicate a reduction in symptoms or regression to the mean. Using both categorical and continuous approaches to classification and consistent with prior literature documenting the stability of these constructs, behavioral and emotional risk screening classifications were largely stable across time. Overall, these results bode well for schools engaged in less frequent screenings, as the likelihood of students moving from no risk to significant risk is minimal, and those with elevated risk tended to return to normal risk.
Latent profile analyses provided unique insight into the most likely stable patterns of self-reported emotional and behavioral risk that occurred within this longitudinal sample. Results suggested a three-class solution with one class of students with minimal to no risk across time, another class with average to high average risk across time, and another class with significant risk across time. Self-reported screening scores were consistently stable with the majority of students being easily classified into one of these categories. Results suggest that once a student is identified as having low risk, normal risk, or high risk, then they are likely to remain so across a 4-year time period. Of greatest concern are the students identified as consistently at high risk. The longitudinal outcomes for students identified as at-risk are unknown, although previous research suggests that BESS Student screening scores are predictive of a variety of negative behavioral outcomes, such as office disciplinary referrals, suspensions (Chin, Dowdy, & Quirk, 2013), substance use, fighting, and suicidal ideation (Dowdy, Furlong, & Sharkey, 2013).
These results support current screening practice recommendations to attend to the early identification and early intervention of behavioral and emotional problems (Glover & Albers, 2007). Schools may want to consider adopting a multiple-gating screening approach (e.g., Severson et al., 2007) to be able to efficiently follow-up with the students identified as at risk, particularly as they are likely to remain at risk. Students who were initially screened as having low risk can subsequently be rescreened less often, whereas those with elevated or extremely elevated risk can be offered intervention services and be rescreened more frequently to help determine the effectiveness of the interventions. As emotional and behavioral screening scores are likely to remain stable across time, there is a critical need for early intervention that can decrease the likelihood of negative associated outcomes.
Implications for Practice
Screening, regardless of when or how frequently it is conducted, is only helpful if effective services are provided to the youth who are in need of additional services (Shirk & Jungbluth, 2008). Screening for emotional and behavioral risk is likely to become a mainstay within the educational system, particularly as results are easily incorporated into data management systems (e.g., AIMSweb) already in use for the assessment and progress monitoring of academic and behavior problems. Therefore, it is incumbent on professionals to not only determine the practicalities associated with screening but also those associated with follow-up preventive and early interventions. Although additional research is needed to systematically study the frequency with which screenings should be repeated and the age at which screenings should commence (O’Connell, Boat, & Warner, 2009), this study provides preliminary evidence that self-reported screening scores are at least moderately stable over a 4-year period.
Implications are straightforward for the students who were classified into the group with significant risk across time. Most critically, the results of this study suggest that those students identified as demonstrating significant risk are most likely to remain at-risk over time; therefore, intervention efforts to address such risk are even more imperative in order to help mitigate the exacerbation of such problems. Screening at an earlier age, prior to more stable patterns emerging, may allow for early intervention efforts to help place students on a more adaptive trajectory. Some research has pointed to this need for early screening, starting at the preschool level (Dowdy, Chin, & Quirk, 2013), whereas other research suggests that screening at the age of 5 years may detect few cases of later significant mental illness and that it may be more useful to wait until adolescence to screen for depression or anxiety (Najman et al., 2007). Future research can determine when patterns of risk begin to stabilize, and relatedly when it may be optimal to begin screening so that intervention efforts have a higher likelihood of being successful.
Students initially classified into either the group with minimal to no risk or the group with average to high average risk are likely to remain within the Normal range, or below the threshold needed for additional assessment or intervention activities. However, critical events have the potential to change a student’s emotional or behavioral status rapidly and school personnel need to be aware of this and to avoid basing service provision solely on the results of screenings, regardless of when the screening was conducted. Although additional research is needed prior to concluding how often screenings need to be conducted, these results seem to contradict screening practices in which data are collected at multiple time points throughout an academic year (Ennis et al., 2012). Findings show that screening profiles, regardless of level of risk, remain consistent over time, suggesting that it may be unnecessary to screen multiple times throughout the year. These results have implications for the efficiency of education services, as the costs of screening at multiple time points within an academic year are likely to greatly outweigh the potential benefits of discovering new cases of behavioral or emotional risk.
Limitations and Future Directions
Although we recognize it is not feasible within the scope of this study to provide detailed recommendations regarding the exact timing and intervals between screenings, this investigation was intended to provide broad information regarding the likelihood of students’ ratings and their resulting categorical risk status and screening profile to remain consistent over the course of several years. However, there are limitations to this study that deserve mention so that readers can use discretion when interpreting the findings.
One significant limitation of this study is the relatively small sample size available. With larger sample sizes and more representation in the elevated risk categories, additional classes of longitudinal risk may be identified and movement across categories can be examined more closely. Furthermore, studies with larger sample sizes would permit an examination of those students who do shift among risk categories, and potential predictors of such changes over time. Additionally, the sample largely consisted of Latino/a and White students and was limited to students in Grades 8 to 11, with data collected at 1-year intervals through self-report.
The reliance on self-report data alone is another limitation of the study, particularly as adolescents may not accurately describe externalizing behaviors (Smith, 2007). This limitation may have affected the findings of similar stability coefficients across externalizing and internalizing problems. Unfortunately, in this study there were not additional measures, such as direct observation or teacher report, to which the stability of the self-report BESS Student screening results could be compared. Future research should compare the stability of different types of problems with multiple approaches and using multiple informants. Although current data evidenced self-reported stability across time, adolescents may be consistently guarded or inaccurate in their self-reports, suggesting that triangulation of data is needed for clinical purposes. Use of multiple informants may be particularly needed, as some research has shown interrater reliability ratings of externalizing problems to differ across the school year with agreement increasing over the second half of the school year (Evans, Allen, Moore, & Strauss, 2005). Future research is needed with a variety of different samples, different informants, and with different and more frequent retest intervals. Ultimately, to determine when and how often screenings are to occur, longitudinal studies examining screening and various outcomes of interest at varying time points are needed. Similarly, cost-effectiveness studies that explicitly examine the balance of any potential benefits that may occur from screening with any potential disadvantages are needed (Chatterji, Caffray, Crowe, Freeman, & Jensen, 2004).
Stability can be affected by both changes in the individual over time, as well as error in measurement (American Educational Research Association, American Psychological Association, & National Council on Measurement Education, 1999), and it is unknown if some individuals received interventions directly targeted to reduce their behavioral and emotional risk during the course of this study. In general, universal prevention programs were offered at junior high and high schools as part of the SSHS project, and there was greater overall access to mental health services at the schools. However, BESS Student results did not directly inform follow-up services, which is not consistent with best practice recommendations. Also, if a student received mental health services, that may have changed their risk level status, and we do not know which students may have received interventions, due to confidentiality issues. It is possible that developmental changes contributed to change in risk across time or that the screening instrument is more sensitive to change at certain ages (Watkins & Smith, 2013). Future research is needed to determine if risk status changes as a result of intervening factors, providing additional utility for screening measures to be used in progress monitoring and intervention evaluation.
Results also need to be interpreted in light of attrition analyses findings, which indicated that students with higher levels of behavioral and emotional risk were more likely to drop out of the study. Although this is in line with expectations that behavioral and emotional risk is associated with poorer school behavioral outcomes, results may differ for students with more significant symptoms at an initial screening time point.
Future research is necessary to determine how much of the variance in change in risk status over time is predictable based on known variables. In addition to individual variables such as participation in intervention as mentioned above, other individual-level variables, such as age, may be predictive of malleability over time. As this study focused on adolescents, and more specifically adolescent self-report, more research is needed to determine how frequently screenings should be conducted among younger children, and whether stability varies by informant or type of assessment (e.g., questionnaire vs. observation). Furthermore, at the school level, contextual risk and protective factors may help predict change in risk status over time. Research is needed to identify school-level variables and environmental risk factors that could help explain movement in risk over time.
Investigations examining the students who do change categories are needed to further determine the quantitative nature of their changes, as well as the qualitative differences among students who do and do not change across time. Through further examination into students who do and do not change, we may be able to discern when a student needs additional follow-up screenings or additional assessment. Although the practice of screening for behavioral and emotional risk is still highly advised based on decades of research supporting the value of preventative and early intervention activities (Glover & Albers, 2007), this study highlights the need for additional research to make data-based decisions regarding when and how often to conduct screenings.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported in part through the evaluation of a Safe Schools Healthy Students grant (Principal Investigators: Drs. Michael Furlong, Jill Sharkey, Erika Felix, and Erin Dowdy).
