Abstract
The present study investigated evidence of the construct validity of scores from the Behavioral and Emotional Rating Scale (BERS-3), which is a multi-informant assessment designed to measure the behavioral and emotional strengths of school-aged youth. The purpose of this research was to evaluate the degree to which BERS-3 scores differed between students with school-identified emotional disturbance and students without disabilities. Two nationally representative samples were used in this study: (a) 1,575 students rated by teachers and (b) 793 youth who provided self-ratings. The results of multivariate multiple regression analyses supported the primary hypothesis that students with emotional disturbance would have lower scores on each of the five BERS-3 subscale scores compared to peers without disabilities. This finding held for both samples; however, differences between students with emotional disturbance and the peers without disabilities were substantially smaller for the youth self-ratings compared to teacher ratings. Implications for practice and directions for future research are also discussed.
Keywords
Assessment instruments are used to understand children’s functioning and inform decision-making. These educational, mental health, and social service decisions involve collecting significant amounts of data across raters on the child, diagnosing the child and placing the child into specialized programs and services, and then evaluating the outcomes of these programs. While several assessment instruments are available to collect data about a child for decision-making purposes, these test instruments often focus on identifying a child’s deficits, problems, and pathologies. Although these instruments meet acceptable levels of reliability and validity and are useful in identifying children who can benefit from interventions and supports, they are less helpful in providing a complete picture of the child and in designing a comprehensive individualized education program (IEP) or treatment plan (e.g., Chatzinikolaou, 2015; Climie & Henley, 2016). Over the past few decades, there has been a call for measuring the emotional and behavioral strengths or competence of children rather than a deficit orientation, which has been referred to as strength-based assessment (Epstein, 2004).
The idea of measuring the strengths, competencies, assets, and resources of a child is congruent with a focus on promoting children’s well-being. The United Nations Convention on the Rights of the Child (United Nations General Assembly, 1989) identified that it is every child’s right to reach an acceptable level of development, education, and overall quality of life. The concept of child well-being and the Rights of the Child have gained widespread acceptance from professional (e.g., American Psychological Association, Division 16; International School Psychology Association) and international governing groups (e.g., the European Commission, Office for Economic Cooperation and Development). Although the importance of child well-being has been recognized, the concept has been elusive to define (Ben-Arieh, 2008). However, one area of agreement is that child well-being is not merely the absence of limitations or problems but is the presence of positive attributes such as personal strengths, competencies, resources, supports, and assets (Ben-Arieh, 2008). The use of strength-based assessment (SBA) during treatment may contribute to an increase in youth functioning and better parent satisfaction. Specifically, Cox (2006) reported that child functioning outcomes were significantly better for children who received a strength-based oriented assessment as opposed to a deficit-based assessment. Additionally, parents were more satisfied with the strength-based assessment model and the resulting treatment plan. In general, highlighting child, environmental, and contextual strengths can lead to assessment results that are more meaningful and persuasive to parents. In so doing, the likelihood that assessment results will be accepted and recommendations implemented by parents is greatly increased (Mastoras et al., 2011).
Driven by over two decades of a research on developmental assets theory, a focus on children’s emotional and behavioral strengths and child well-being has generated much interest among professionals. Advocates of strength-based assessment argue that even the most troubled and challenging children and youth possess strengths that can be tapped in the service of adaptive approaches to treatment or intervention (Buckley & Epstein, 2004; Cox, 2006; Woodland et al., 2011). There are numerous benefits to the use of SBA including a more well-rounded representation, balanced and preventive view of the child while also acknowledging both ecological and contextual variables, and their contribution to strengths and challenges. Indeed, awareness of strengths and contextual variables can be leveraged to inform recommendations and interventions ultimately improving both treatment acceptability and likely utility (Climie & Henley, 2016). Further, assessing a child’s personal or environmental strengths can provide information on assets that can be tapped for use in the intervention process (Snyder et al., 2006).
Along with this professional interest in operationalizing and measuring children’s strengths and well-being, a number of tests have been developed to assess these attributes (e.g., Merrell, 2011). Perhaps one of the most widely used strength-based instruments is the Behavioral and Emotional Rating Scale (BERS; Epstein, 2004; Epstein & Sharma, 1998), which is a standardized, norm-referenced instrument that measures youths’ emotional and behavioral strengths. The BERS is an assessment system that includes three rating scales: Teacher Rating Scale, Youth Rating Scale, and Parent Rating Scale. The three versions contain the same 52 items, although there are minor wording modifications in items to accurately reflect the perspective of the informant. The BERS provides five subscales of emotional and behavioral strengths (i.e., Interpersonal Strength, Family Involvement, Intrapersonal Strength, School Functioning, and Affective Strength) and a Total Strength Index. Although the reliability and validity of the original BERS as a strengths-based measure was well established, two limitations prompted a 2001–2002 re-norming with a large representative sample of parents/caregivers and children and youth. Specifically, the original BERS did not differentiate between parents and teachers nor did it permit adolescents to report on perceptions of their own strengths and limitations (Buckley & Epstein, 2004). The re-norming of the BERS resulted in the BERS-2: Parent Rating Scale (Epstein, 2004), the BERS-2 Youth Rating Scale (Epstein, 2004), and the re-standardized BERS 2: Teacher Rating Scale (Epstein, 2004).
An important criterion in evaluating a test is validity, which refers to whether or not a test actually measures what it purports to measure (Messick, 1995). A number of studies indicated that the original BERS yields scores that are psychometrically sound (e.g., Epstein & Sharma, 1998; Epstein et al., 1999; Epstein et al., 2002). However, when a test is re-normed, professional organizations, practitioners, and researchers recommend that the validity and other aspects of the test’s psychometrics be re-established (American Educational Research Association, 2014). The process of validating an assessment instrument involves collecting a significant amount of data to provide a framework for evaluating an instrument’s scores (American Educational Research Association, 2014). A series of investigations of the BERS scores have provided validity evidence based on internal structure, convergent relationships, and test-criterion relationships (see Epstein & Sharma, 1998; Epstein et al., 1999; Epstein et al., 2002; Sointu, et al., 2014; Lambert et al., 2019).
Construct validity is a specific type of validity, which refers to the degree to which underlying constructs of a test can be identified and the degree to which these constructs reflect the theoretical foundation on which the test instrument is based (Nunnally & Bernstein, 1994). As the BERS assesses emotional and behavioral strengths of children, the test scores should discriminate between children who are known to evidence greater emotional and behavioral strengths and students who are known to evidence fewer of those strengths. One particular group of students who are known to evidence fewer emotional and behavioral strength are students with emotional disturbance (ED). Students with ED exhibit one or more of the following characteristics over an extended period and to a marked degree that adversely affects their educational performance: (a) an inability to learn that cannot be explained by intellectual, sensory, or health factors; (b) an inability to build or maintain satisfactory interpersonal relationships with peers and teachers; (c) inappropriate types of behavior or feelings under normal circumstances; (d) a general pervasive mood of unhappiness or depression; and (e) a tendency to develop physical symptoms or fears associated with personal or school problems. Previous research with the BERS suggests that scores differentiate between youth with and without ED or emotional and behavioral disorders (Uhing et al., 2005). However, the validity of a test cannot be determined by a single research study, but by several studies that examines the relationship between the test and the behavior it purports to measure (Messick, 1995). Thus, one purpose of the present research was to continue to evaluate the psychometric functioning of the BERS in relation to its construct validity.
Like the original BERS, studies support the psychometric integrity of scores from the BERS-2. Evidence for the convergent validity of scores has been established in several studies (Epstein, 2004; Lambert et al., 2019). The BERS-2 Youth Rating Scale subscales and Strength Index scores were found to possess positive associations with scores from the Social Skills Rating System-Student Form composite (Secondary Level, Grades 7–12) and negative correlations with scores from the Problem scales of Achenbach’s Youth Self-Report—both support evidence of convergent validity. Convergent validity of scores from the BERS-2 Teacher Rating Scale and the Achenbach Teacher’s Report Form (TRF: Achenbach, 1991) has also been found to be strong between the BERS-2 and the TRF (Benner et al., 2008). Strong test–retest reliability has also been established for the BERS-2 Youth Rating Scale with reliabilities for all scales exceeding .80 over a 1 week period. Cross-informant agreement between parents and youth has also been established to be moderate to high ranging from .50 to .63 for the BERS-2 (Synhorst et al., 2005). Studies have also shown stability over time of parent ratings over a 1-year (r = .78) and 2-year period (r = .71) (Lambert et al., 2015), and student behavioral and emotional strengths have demonstrated a positive relationship with student–teacher relationships and academic achievement (Sointu et al., 2017). In all, results from multiple studies support the use of the BERS-2 in assessment of socio-behavioral functioning of children and youth.
Evidence indicates there may be differences in youth’s emotional and behavioral functioning as a function of gender and age. There is some evidence of gender differences in protective factors in the general population (Hartman et al., 2009) as well as with youth with an emotional or behavioral disorder (Novak et al., 2020). Within the strengths-based assessment literature, the research on gender differences is mixed, with some research findings that girls may exhibit greater strengths than boys (Romer et al., 2011; LeBuffe et al., 2018) and others finding no differences as a function of gender (Epstein et al., 2002; Gilman & Huebner, 2006). Regarding age, prior research suggests no differences as a function of enrollment in elementary or secondary grades (Epstein et al., 2002). Determining whether there are gender or age differences in ratings of youth’s emotional and behavioral strengths can inform the assessment and treatment planning process. For example, if younger youth tend to have higher levels of strengths compared with older youth in the area of family involvement, this information on family strengths can be used to enhance family participation in the intervention or treatment plan.
The primary purpose of this study was to re-establish evidence of the content validity of scores from the BERS-3 Teacher Rating Scale (TRS) and the Youth Rating Scale (YRS) by evaluating the extent to which behavioral and emotional strengths differed between students with and without emotional disturbance (ED). Consistent with prior research (e.g., Uhing et al., 2005), we hypothesized that the TRS and YRS scores of students with ED would demonstrate significantly lower strengths than students without ED and that the effects would be large in magnitude. Secondary research purposes involved exploring the relationships between strengths and gender and age: specifically, (a) the extent to which behavioral and emotional strengths differed across gender and age-groups and (b) the extent to which the differences in strengths between students with and without ED were moderated by gender and age. We hypothesized that student strength scores would not differ by these demographic variables; however, this hypothesis should be considered exploratory given the relative lack of research on gender and age differences. We did not postulate an a priori hypothesis regarding the direction or magnitude of the moderated effects.
Method
Participants
This study included two samples of students who were rated by their teacher (on the Teacher Rating Scale) and students who completed the Youth Rating Scale. Both samples of students were drawn from the norming study for the Behavioral and Emotional Rating Scale–3 (BERS-3; Epstein et al., 2021). Participants in the teacher-rated sample included 1,575 students. Slightly less than 16% of the sample (n = 246) represent students school identified with emotional disturbance (ED) and were receiving special education services. The remaining 84% of the sample (n = 1,329) were students without any identified disabilities. Students ranged in age from 5 to 18 years with a mean age of 11.45 years (SD = 3.67). The sample was split fairly evenly split between male (53%; n = 836) and female students (47%; n = 739). The majority of students identified as white/non-Hispanic (51%; n = 809) with smaller proportions of Black or African-American/non-Hispanic (20%; n = 314) and Hispanic or Latinx students (17%; n = 269). Just over 37% of students (n = 583) received free or reduced lunch; however, these data were missing for 20% of the sample (n = 322). Fewer than 2% of students (n = 25) spent more than half of the school day in special education settings. Nearly 79% of students were rated by female teachers (n = 1,209), and 87% were rated by a teacher who identified as white/non-Hispanic (n = 1,366).
Participants in the self-rated sample included 793 students. Just over 6% of the sample represented students with ED and were receiving special education services (n = 48) while the other 94% of the sample were students without disabilities (n = 745). Students ranged in age from 11 to 18 years with a mean age of 14.01 years (SD = 2.02). The sample consisted of slightly more female students (52.7%; n = 418) than male students (47.3%; n = 375). The majority of students identified as white/non-Hispanic (55.7%; n = 442) with smaller proportions of Black/non-Hispanic (11.5%; n = 91) and Hispanic students (23.5%; n = 186). Nearly 38% of students (n = 299) received free or reduced lunch; however, these data were missing for 5.8% of the sample (n = 46). Just over 1% of students (n = 10) spent more than half of the school day in special education settings.
Data Collection
Before data were collected, three university Internal Review Boards (University of Nebraska-Lincoln, University of Northern Colorado, and Elon University) approved recruitment and data collection protocols. Data were collected from Fall 2015 through Spring 2018 as part of the re-norming of the BERS-3 in the following manner. First, the authors of the BERS-3 contacted local district administrators and university professionals across the four United States geographic regions (North, Midwest, South, and West) to identify teachers who might volunteer to participate in data collection. Teachers were contacted by one of the authors either by email, mail, or telephone and asked to participate in the norming process. Teachers who agreed to participate were asked to complete the instrument on all of their students or to select an unbiased sample of their students using a simple procedure. Specifically, raters were given the following instructions to ensure an unbiased selection process. “First, decide how many students you wish to rate using the TRS. Then, start at the top or bottom of your class roster and rate every student. Do not skip any student unless you have known this student for less than 2 months. Stop selecting and rating students when you have reached the number of students you wished to rate.” Finally, teachers who agreed to collect completed YRS forms then contacted parents detailing the purpose of the study, consent procedures, and a consent document they could sign and return to school. For students whose parents provided consent, the teachers explained to their students the data collection process and sought their assent. Returned TRS and YRS forms were reviewed by the research team project manager, and any rating scales with missing demographic data (i.e., age in years, geographic region indicator, gender, race, Hispanic/Latinx status, or exceptionality status) or item responses were excluded from the study.
Measure
Like its predecessor, the BERS-2, the BERS-3 (Epstein et al., 2021) consists of three forms—the Teacher Rating Scale (TRS), the Parent Rating Scale (PRS), and the Youth Rating Scale (YRS), although only data from the TRS and YRS were used in this study. Each form consists of 52 parallel items that a teacher, parent, and student rate to assess the emotional and behavioral strengths of the student. The items address five areas of behavioral and emotional strength that measure the positive emotions, behaviors, and aspects of an individual’s life (Epstein et al., 2021). Thus, the BERS-3 consists of five core subscales (Interpersonal Strength [IS], Family Involvement [FI], Intrapersonal Strength [IaS], School Functioning [SF], and Affective Strength [AS]) for ages 5 through 18.
The Interpersonal Strength subscale measures a student’s ability to control his or her emotions or behaviors in social situations (e.g., accepts “no” for an answer; reacts to disappointments in a calm manner). The Family Involvement subscale measures a student’s participation in and relationship with his or her family (e.g., participates in family activities; interacts positively with siblings). The Intrapersonal Strength subscale measures in a broad sense a student’s outlook on his or her competence and accomplishments (e.g., is self-confident; is enthusiastic about life). The School Functioning subscale focuses on the student’s competence in school and classroom tasks (e.g., pays attention in class; completes school tasks on time). The Affective Strength subscale assesses a student’s ability to accept affection from others and express feelings toward others (e.g., accepts a hug; asks for help).
The TRS is completed by an educator who is familiar with the student, and the YRS is completed by the student. Each item is rated on a 4-point Likert scale. TRS items are rated 3 = “very much like the student,” 2 = “like the student,” 1 = “not much like the student,” and 0 = “not at all like the student”; YRS items are rated 3 = “very much like you,” 2 = “like you,” 1 = “not much like you,” and 0 = “not at all like you.” Each of the subscales is summed to obtain a raw score which is transformed to a scaled score, with a mean of 10 and a standard deviation of 3. Standard scores were computed based on a single distribution of raw scores for the entire normative sample rather than based on distributions of raw scores for specific subgroups (e.g., men, women, children with ED, and children without disabilities). The developers chose to compute one set of standard scores (and one set of normative standards) because initial analyses of differential item functioning (DIF) indicated invariance across subgroups of children (Epstein et al., 2021).
Data Analysis
The primary focus of the analysis was to evaluate the extent to which BERS-3 subscale scores differed between students with and without ED (i.e., the main effect of ED). To this end, STATA v13 (Statacorp, 2013) was used to analyze data within a multivariate multiple regression framework where all five BERS-3 subscale scores were the dependent variables. Predictors in the models included a dummy-coded variable representing a contrast between students with ED and students without ED (with students with ED as the focal group), dummy-coded student age (with elementary-aged students as the focal group), and dummy-coded student gender (with male as the focal group). Age and gender were primarily included to account for differences in demographics between students with and without ED as well as to explore secondary research questions.
After testing the main effects of ED, age, and gender, two interactions were added to the regression model, ED*Age and ED*Gender, to evaluate whether the effect of ED was dependent on age or gender (i.e., moderated by age or gender). Statistically significant interactions were then probed to evaluate the nature of the interaction by computing simple effects for ED at each level of the moderator (e.g., age or gender) (Aiken & West, 1991). A significant interaction indicates that the simple effects of ED differ across levels of the moderator. A simple effect of ED represents the mean difference between students with ED and their peers without disabilities at a single level of the moderator (e.g., the effect of ED for female students). The statistical significance of simple effects was obtained using the contrast command in STATA.
Because the distributions of BERS-3 subscale scaled scores (and residuals) were non-normal, the standard errors for model parameters were computed using non-parametric bootstrapping based on 1,000 bootstrapped replications. Cohen’s d effect sizes (1988) were computed for the main effect and simple effect comparisons between the groups. Cohen’s d statistics represent the mean difference between two groups in standard deviations units. Cohen’s d estimates were computed using the model-adjusted means from the regression analysis (e.g., controlling for age and gender) and the unadjusted variances as recommended by the What Works Clearinghouse (2020). We characterized each effect size estimate according to Cohen’s suggested ranges for the d effect size statistic (Cohen, 1988); therefore, d values ranging 0.00–0.19 were designated as trivial, 0.20–0.49 were small, 0.50–0.79 were medium, and ≥0.80 were large. For significant interactions, Cohen’s f 2 was computed as a measure of effect size, which indicates the proportion of explained variance attributable to the interaction. In general, Cohen (1988) suggested that f 2 values of .02, .15, and .35 indicate small, medium, and large effects, respectively; however, these guidelines were primarily developed to interpret the magnitude of main effects rather than interactions. A review of the applied psychological research revealed that the mean interaction effect size was f 2 = .009 and the median effect size was f 2 = .002 (Aguinis et al., 2005), so we interpreted the interaction effect sizes within this context rather than Cohen’s general guidelines.
A total of 15 main effect comparisons per sample were interpreted—three comparisons per subscale score (i.e., ED, age, and gender). If the significance level was set at a .05 per-test level, there would be a 54% chance of making at least one Type I error across the set of comparisons and a 17% chance of making two errors per sample. To account for the Type I error rate inflation caused by multiple comparisons and maintain a nominal significance level of .05 for the entire set of main effect comparisons per sample, we adopted a conservative per-test significance level of .0034. A total of 10 interactions were interpreted per sample—two per subscale score (i.e., ED*Age and ED*Gender). Because statistical tests for interactions are significantly less powerful than tests for main effects (at a given sample size), we did not adjust for multiple comparisons and set the per-test significance level at .05.
Results
Teacher Ratings
Unadjusted Means and Standard Deviations by Group for Teacher Rating Scale Scores.
Zero-Order Correlations for Teacher Rating Scale Scores.
Results from Multivariate Multiple Regression Analysis for TRS Scores.
Youth Ratings
Unadjusted Means and Standard Deviations by Group for Youth Rating Scale Scores.
Zero-Order Correlations for Youth Rating Scale Scores.
Results from Multivariate Multiple Regression Analysis for YRS Scores.
Note. The significance level was set at p < .0034 for main effects and p < .05 for interactions. YRS = Youth Rating Scale. ED = emotional disturbance.
The interaction between ED*Elementary was statistically significant for the School Functioning subscale score (bED*Elementary = 2.19, p = .017, f 2 = .009) indicating that the simple effect of ED differed for students in elementary school compared to students in secondary school. The effect size for the interaction suggests a fairly large effect. The simple effect of ED for students in elementary school was non-significant and positive (bED/Elementary = 1.32, p > .05, d = 0.59) indicating that students with ED rated themselves with slightly higher scores than their peers without disabilities. On the other hand, the simple effect of ED for students in secondary school was significant and negative (bED/Secondary = −1.19, p = .007, d = −0.40) indicating that students with ED has rated themselves with lower scores than their peers without disabilities. This interaction is plotted in Figure 1. Plot of ED*Age interaction for School Functioning.
The interaction between ED*Female was statistically significant for the Intrapersonal, Family, and Affective subscale scores indicating that the simple effect for ED differed significantly across female and male students. For the Intrapersonal Strengths subscale score, the interaction was negative (bED*Female = −4.14, p < .001, f 2 = .023) indicating that the simple effect of ED was larger for female students compared to male students. The effect size for the interaction suggests a large effect. The simple effect of ED for female students was statistically significant and negative (bED/Female = −4.91, p < .001, d = −1.70) indicating that female students with ED rated themselves with significantly lower strengths compared to their male peers without disabilities. For male students, the simple effect of ED was not significant (bED/Male = −0.76, p > .05, d = −0.26) indicating that male students with ED did not rate their strengths differently than male peers without disabilities. This interaction is plotted in Figure 2. Plot of ED*Gender interaction for Intrapersonal Strengths.
For the Family Involvement subscale score, the interaction was negative (bED*Gender = −3.12, p = .010, f 2 = .012) indicating that the simple effect of ED was larger for female students compared to male students. The effect size for the interaction suggests a medium to large effect. The simple effect of ED for female students was statistically significant and negative (bED/Female = −4.17, p = .024, d = −1.29) indicating that female students with ED rated themselves with significantly lower strengths compared to their female peers without disabilities. For male students, the simple effect of ED was not significant (bED/Male = −0.89, p > .05, d = −0.29) indicating that male students with ED did not rate their strengths differently than male peers without disabilities.
For the Affective Strengths subscale score, the interaction was negative (bED*Gender = −2.05, p = .040, f 2 = .005) indicating that the simple effect of ED was larger for female students compared to male students. The effect size for the interaction suggests a small to medium effect. The simple effect of ED for female students was statistically significant and negative (bED/Female = −3.13, p = .020, d = −1.07) indicating that female students with ED rated themselves with significantly lower strengths compared to their female peers without disabilities. For male students, the simple effect of ED was not significant (bED/Male = −1.22, p > .05, d = −0.40) indicating that male students with ED did not rate their strengths differently than male peers without disabilities.
Discussion
Research on previous editions of the BERS scores has suggested acceptable validity evidence based on internal structure and convergent relationships (Epstein & Sharma, 1998; Epstein, 2004). The purpose of the present investigation was to extend the previous validity studies to the BERS-3 scores by investigating their construct validity. Overall, the findings supported the primary hypothesis. Specifically, the TRS scores of school-identified students with ED significantly differed from those of students without disabilities. Moreover, these differences were present across each of the subscales, were in the predicted direction, and were of a large magnitude. However, for YRS scores, the findings were more nuanced.
While the main effects of ED were statistically significant, large in magnitude, and consistent across four of the five subscales of the YRS, the effects of ED were often moderated by gender or age. For example, gender moderated the effect of ED for the Intrapersonal Strengths, Family Involvement, and Affective Strengths subscales in such a way that scores differed between students with and without ED for only female students—male students with ED did not rate their strengths differently than male students without disabilities. These interactions may highlight important differences in how students perceive and rate their own strengths. This may be especially important when comparing youth self-ratings and teacher ratings because the correspondence between the two sets of raters may be dependent of students characteristics such as gender and age.
In terms of the main effects of gender, there were no statistically significant differences between male and female students on teacher or youth self-report rating forms. Compared to older students, however, teachers rated elementary students higher on family involvement and affective strengths while youth self-reported higher than older youth on interpersonal strengths, family involvement, and school functioning. The overall findings of the present study, along with the validity studies from earlier editions of the BERS (Epstein et al., 1999, 2002; Epstein & Sharma, 1998) afford supporting evidence for the validity of the test scores of the BERS-3, in addition to important findings related to the conditionality of scores from the YRS.
The most important finding was that the scores of the children with ED were significantly lower than the test scores of the students without disabilities across each of the five dimensions of emotional and behavioral strengths—Interpersonal Strength, Family Involvement, Intrapersonal Strength, School Functioning, and Affective Strength for teacher ratings and across four of the five subscales for youth self-ratings. The finding of significant differences in strengths provides support for the construct validity of the BERS-3 scores, indicating endorsement for the theoretical model underpinning the BERS development.
Limitations and Future Directions
Several limitations of this study need to be noted. First, the selection of students who were rated was not conducted on a random basis. The professionals who volunteered to participate were contacted by the researchers and asked to participate. Thus, the sample of professionals included only those individuals who volunteered to participate and filled in BERS-3 forms on the children. Thus, this volunteer sample of respondents does not provide data on children with or without disabilities whose teachers or other professionals did not respond and whose responses might be systematically different from those professionals who participated. Second, regression analyses did not include a measure of SES even though families of students with ED tend to over-represent lower SES groups. We collected data on free or reduced lunch (FRL) which would be a sufficient proxy for SES; however, these data were missing for approximately 20% of participants many of which were students with ED. Because the multivariate multiple regression analysis used a pairwise present approach to handling missing data, including FRL would have excluded a large proportion of students with ED.
A third limitation was that minimal data were collected on the demographic background of the professionals who provided the ratings. Future investigators should collect important demographic and professional information on the respondents including gender, age, ethnicity, race, terminal degree, and years working in the field and evaluate how these variables may be related to their ratings. Because data were not collected on the specific employment role of the respondents (e.g., teachers, psychologists, and social worker), it is possible that respondent view of child behavior may have been influenced by their professional role within the setting. Although earlier cross informant research has demonstrated that individuals with similar roles (e.g., parent to parent; teacher to teacher) are in reasonably high agreement (Achenbach et al., 1987), future investigators need to determine the cross-informant agreement between respondents of the BERS-3. Fourth, the professionals who completed the BERS-3 ratings forms were aware which children were and were not identified with ED. Prior knowledge of a child’s status may have influenced bias in that individuals’ child ratings leading to greater discrimination between the identified and non-identified groups of children. Future researchers should consider studying the predictive validity of BERS-3 scores for a sample of children who have not already been identified as experiencing emotional and behavioral challenges. Finally, the number of students with ED who completed the YRS numbered less than 50 students and was much smaller than the sample of students without disabilities. Other investigators need to replicate the study with a greater number of students with ED.
The current investigation along with the research on previous editions of the BERS provided support for several aspects of the validity of the scores; yet there exist other directions for future investigation. For example, researchers have not investigated whether the BERS-3 scores are biased in assessing the emotional and behavioral strengths of children of different racial and ethnic backgrounds. The lack of bias or fairness of scores from children from diverse backgrounds including Black or African-American, Hispanic or Latinx, Asian, Native American, and multi-racial children should be investigated. This point is particularly relevant as the US child population is estimated to become significantly more racially and ethnically diverse in the coming years (Hussar & Bailey, 2016). To this end, test developers need to demonstrate that the test instruments are acceptable and fair for use with children from diverse, multicultural backgrounds. Finally, future investigators need to evaluate the social validity of the BERS-3 by asking professionals about the instrument’s feasibility in the treatment planning process. Tests with high social validity tend to have greater consumer acceptability, higher treatment expectations, and greater treatment fidelity (Donovan & Nickerson, 2007).
Implications
This study has several implications for practice. The scale of the differences in test scores between children with and without ED underscores that significant behavioral differences are present between these student populations, which suggest several teaching, professional preparation, and research implications. For example, when used with other assessment data, BERS information can assist parents, teachers, psychologists, and other professionals to identify behaviors to be strengthened, to set goals for IEPs or treatment plans, and to build on current strengths. Indeed, BERS information can be used in case conceptualization and subsequently tapped to guide intervention goals. As an example, strengths on the Family Involvement subscale that assess a student’s participation in and relationship with his or her family can be effectively used to shore up family presence and involvement in support of a struggling adolescent’s coping resources and strategies. Also, teacher training programs need to prepare individuals not to merely identify deficits and problems but with knowledge and skills in strength-based assessment in order to focus decisions more positively on strategies to foster child strengths while working to bolster areas where children are not as strong. Finally, researchers involved in studying the behaviors of students with ED should focus attention on the emotional and behavioral strengths of these children and how strengths may mitigate problem behavior. For example, strength-based assessment data from the BERS could be used to operationalize constructs such as resilience and protective factors which have been found to enhance positive life outcomes of children and youth.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
