Abstract
The purpose of this study was to examine the influence of primary performance area, education level, and performance quality on pre-service music teachers’ evaluations of middle school string orchestra performances. Participants (N = 78) were pre-service band, choral, and orchestra teachers who self-reported their academic status as lower (n = 39) and upper (n = 39) classmen. Participants assigned ratings to interpretation-musicianship, dynamics, balance/blend, and other factors on a 7-point Likert-type scale with criteria-specific descriptors. Repeated-measures ANOVA tests revealed that participants were able to distinguish between both good and poor performances. Upper classmen pre-service music teachers assigned more favorable ratings to interpretation-musicianship and balance/blend than lower classmen. Pre-service choral teachers gave less favorable ratings than pre-service orchestra teachers for interpretation-musicianship and balance/blend. Descriptive analysis revealed that upper classmen pre-service music teachers assigned more favorable ratings than lower classmen. For all evaluation statements, pre-service choral teachers gave the least favorable ratings and pre-service orchestra teachers assigned the most favorable.
Performance assessments serve as a crucial component of the music education process (Austin, 1990) and some believe adjudicated events are an important curricular element (Garman, Boyle, & DeCarbo, 1991; Rohrer, 2002). The importance of adjudicated festivals to the teaching profession may increase with state legislatures requiring school districts to include student achievement as a component of teacher evaluations (Hassel & Hassel, 2009; Lee, 2009; Williams, 2009). As a result of increased legislation requiring numerical data on teachers, administrators may believe that festival ratings allow for the standardized assessment of students, ensembles, and music educators (NAfME, 2011). In addition to impacting teacher evaluations, low festival ratings appear to negatively impact music programs. Batey (2002) found that poor festival ratings decreased student retention and increased students’ negative attitudes toward their teacher. With a potential growing emphasis on festival results in teacher evaluations and student enrollment in music programs, music educators need to understand who assesses their ensembles at adjudicated events and how performances are evaluated.
As a result of the increased amount of solo, small ensemble, and large ensemble adjudicated events, festival organizers indicated difficultly hiring adjudicators familiar with string pedagogy and repertoire (Barnes & McCashin, 2005). Their survey also revealed string teachers’ concerns about the lack of adjudicators with string experience at orchestra festivals. Findings from prior investigations have supported string teachers’ concerns. High levels of internal consistency were found for adjudication panels that employed judges with knowledge of the performed repertoire and expertise in the field of the performing ensemble (Kinney, 2009). Repp (1996) found that musicians who recently studied the repertoire detected more performance errors than participants unfamiliar with the music. Others concluded that adjudicators’ accuracy when assigning ratings increased with experience (Ekholm, 1997; Wapnick & Ekholm, 1997). Findings from Garman et al. (1991) revealed that inexperienced orchestra adjudicators with backgrounds in composition and private studio teaching demonstrated a lower level of inter-rater reliability than judges with a music education background. To increase the reliability of adjudicators’ ratings at festival, numerous associations that organize adjudicated events began training judges. The effect of training on the reliability of adjudicators’ ratings at festivals has been inconsistent (Boeckman, 2002; Brakel, 2006; Fiske, 1977; Hunter & Russ, 1996; Winter, 1993).
Past investigations on the influence of adjudicators’ primary performance area have provided conflicting results. Some researchers found that adjudicators’ primary performance areas affected overall music performance evaluations (Fiske, 1977; Roberts, 1975; Thompson & Williamon, 2003; Wapnick et al., 2005), while others indicated that primary performance area did not influence evaluations (Fiske, 1975; Geringer, Allen, MacLeod, & Scott, 2009; Hewitt, 2007; Hewitt & Smith, 2004; Massel, 1978; Mills, 1987; Pope, 2012a; Siddell-Strebel, 2007; Simons, 2005; Winter, 1993). Regardless of performing experience on string instruments, no differences were revealed between music majors with no string instrument experience, secondary string instrument experience (non-string players who completed a string methods course), or primary string instrument experience when assigning ratings to general or string specific evaluation statements (Pope, 2012a). Although music majors with different performance backgrounds assigned similar ratings, their self-reported levels of comfort when evaluating string orchestra performances differed. Similar levels of comfort were self-reported by all participants when assigning ratings to the general evaluation statements. However, participants with secondary or no string instrument performing experience indicated lower levels of comfort when giving ratings to string specific aspects of performance. Results revealed that comfort levels of non-string players increased as they gained personal performing experience on string instruments. Conversely, Brakel (2006) found that wind players assigned orchestral performances more favorable ratings than string specialists. Brakel suggested that wind players might have given more favorable ratings than string specialists due to their lack of string pedagogy knowledge.
Results from prior investigations on the influence of adjudicators’ education levels revealed similar evaluations of instrumental performances (Bergee, 1993, 1997, 2003; Geringer et al., 2009; Hewitt & Smith, 2004; Lafferty, 1997; Schleff, 1992; Siddell-Strebel, 2007; Simons, 2005), while others concluded that adjudicators’ education levels influenced their evaluations of instrumental performances (Blom & Poole, 2004; Brittin, Sheldon, & Tian, 2002; Fiske, 1977; Geringer, Madsen, & Dunnigan, 2001; Hewitt, 2002, 2005). Junior high musicians investigated by Hewitt (2002, 2005) gave more favorable ratings than high school musicians and expert adjudicators. A comparison to expert adjudicators’ evaluations revealed that junior high musicians’ ratings for intonation, tempo, interpretation, tone quality, and technique/articulation were less accurate than high school musicians (Hewitt, 2005). In a subsequent study, Hewitt (2007) compared middle school, high school, and university level musicians’ assigned ratings and found that interpretation was the only sub-area to receive similar ratings. However, both middle and high school musicians were more likely to assign less favorable ratings than university level musicians.
Adjudicators have distinguished between different ability levels in both solo and ensemble performances (Byo & Brooks, 1994; Ciorba & Smith, 2009; Geringer & Johnson, 2007; Geringer, Madsen, & Dunnigan, 2001; Johnson & Geringer, 2007; Madsen & Geringer, 1999; Napoles, 2009; Pope, 2012a, 2012b; Schleff, 1992). In Napoles (2009), music majors were able to accurately discriminate between high school and professional choral performances. Music majors (Pope, 2012b) distinguished between professional and all-state orchestral performances by assigning professional orchestras more favorable technique ratings. However, participants did not clearly differentiate between the professional and all-state performances when assessing musicality. In contrast, Doerksen (1999) studied evaluations of concert bands at different performance levels and found that participants did not consistently rate excellent performances more favorably than average performances for intonation, rhythmic precision, and balance/blend.
The current investigation employed an assessment rubric to provide a detailed description of what constitutes the various levels of performance for each evaluation statement. Asmus (1999) identified three specific benefits that rubrics provide over other music performance evaluation tools: (1) provide adjudicators with exact characteristics of performance criteria at all levels; (2) provide specific feedback on the performance to the musicians and teacher; and (3) provide clear indications about specific elements of the performance that need improvement before additional performances. The purpose of the present study was to examine the effect of primary performance area (band, choral, & orchestra), education level (lower & upper classmen), and performance quality (good & poor) on pre-service music teachers’ evaluations of string orchestra performances. Specific questions addressed in this study were: (1) Does primary performance area affect pre-service music teachers’ evaluations of string ensemble performances? (2) Does education level affect pre-service music teachers’ evaluations of string ensemble performances? (3) Are there differences in pre-service music teachers’ evaluations of good and poor quality string ensemble performances?
Method
Participants (N = 78) in this study were undergraduate music education majors from a large university in the southeastern United States. Volunteer participants identified their primary performance areas as band (n = 26), choral (n = 26), and orchestra (n = 26). To assess the possible influence of education level (lower & upper classmen) on the evaluations of string orchestra performances, participants were categorized as lower or upper classmen as a result of their self-reported completion of specific music education courses and the number of completed semesters in their degree program. Lower classmen pre-service band and choral teachers were defined as students with less than four completed semesters of undergraduate study who had not finished the university's Music Education Practicum and String Methods courses. Pre-service band and choral teachers with a minimum of four completed semesters of undergraduate study who finished the university's Music Education Practicum and String Methods courses were defined as upper classmen. At the university where data collection occurred, pre-service orchestra teachers were exempt from the String Methods course and were considered a either lower or upper classmen for this study based on the completion of the Music Education Practicum course and the number of completed semesters in their degree program. Pre-service band, choral, and orchestra teachers were equally represented by lower (n = 39) and upper (n = 39) classmen.
Music Stimuli
Participants were asked to evaluate two middle school string orchestra performances in this study. With the directors consent, performances were video and audio recorded during an adjudicated orchestra festival. To examine possible influences of performance quality (good & poor), the investigator purposefully selected two string ensembles that performed similar repertoire of the same difficulty at contrasting levels of performance. An independent panel of three experienced middle and high school string orchestra teachers unanimously agreed that the two ensemble recordings represented good and poor levels of performance. Each string ensemble performed three selections of music, and the investigator extracted the first two-minutes of each piece with video editing software. The total length of music extracted for each string ensemble was six-minutes. The investigator added a title sequence prior to each string orchestra that identified the performing ensemble (as A or B). In addition, the investigator inserted title sequences that identified the selection of music (1, 2, & 3) performed by each ensemble. Two-minutes of blank video were added at the conclusion of each string ensemble's third selection of music to allow participants adequate time to complete a separate rating form for each group. Directions for competing the study were inserted on the music stimulus DVD prior to the first performance. Each pre-service teacher evaluated a total of 12-minutes of music.
Design and Procedure
Participants (N = 78) in this study evaluated audiovisual recordings of two string ensemble performances using a shortened version of the Indiana State School Music Association's (ISSMA) Official Adjudicator's Comment Sheet: Bands and Orchestras (1999). Participants in this study only assigned ratings to four of the evaluation statements (interpretation/musicianship, dynamics, balance/blend, & other factors) included on the original ISSMA rating form. Brakel (2006) reported acceptable reliability for the Official Adjudicator's Comment Sheet: Bands and Orchestras when used at the Indiana state adjudicated music festival. For each string ensemble, participants rated interpretation/musicianship, dynamics, balance/blend, and other factors using a 7-point Likert-type rating scale with criteria-specific descriptors. The descriptors for each level of performance were: “1” represents outstanding performance in nearly every detail, “1.5” represents some minor flaws, “2” represents frequent minor flaws, “2.5” represents some major flaws, “3” represents frequent major flaws, “3.5” represents continuous major flaws, and “4” represents unacceptable in nearly every detail.
The ISSMA's Official Adjudicator's Comment Sheet: Bands and Orchestras included a description of each evaluation statement that outlined the specific elements of music to be evaluated within the individual statements: Interpretation/Musicianship—consider style, phrasing, tempo, expression, and emotional involvement; Dynamics—consider the appropriate range of dynamic contrast by individuals, sections, and/or full ensemble; Balance/Blend—consider the melodic line, accompanying parts, chord balance, and section/ensemble blend; Other Factors—consider posture, appearance relating to performance, general conduct, and suitable cuts.
All participants evaluated the two middle school string ensembles in a large seminar for music majors. Music stimuli recordings were presented through the built in video projection and sound systems in the classroom. Prior to evaluating the string ensemble performances, participants completed a questionnaire to indicate their experience as a music education major. Participants were also given a brief overview of the study and provided the opportunity to ask procedural questions.
Results
Data collected in this study consisted of participants’ assigned ratings for interpretation/musicianship, dynamics, balance/blend, and other factors during middle school string orchestra performances. Raw data were screened to verify that assumptions of the repeated-measures analysis of variance (ANOVA) were met. Collected data were analyzed with mixed-design repeated-measures ANOVA tests. The analysis included two between-subjects factors (primary performance area & education level) and one within subjects factor (performance quality). A separate repeated-measures ANOVA test was computed for each evaluation statement (interpretation/musicianship, dynamics, balance/blend, & other factors). An alpha level of .01 was used for rejection of the null hypotheses in all statistical tests, and alpha levels were adjusted where appropriate in multiple comparisons. All two- and three-way interactions were non-significant between primary performance area, education level, and performance quality for each evaluation statement. Significant main effects and descriptive data for each variable are discussed below. An overall descriptive analysis revealed that the evaluation statements received the following ratings from least to most favorable: dynamics (M = 2.59, SD = .89), interpretation/musicianship (M = 2.54, SD = .87), balance/blend (M = 2.45, SD = .89), and other factors (M = 2.37, SD = 1.12). Internal reliability for the four evaluation statements (α = .86) was computed by using Cronbach's alpha.
Analysis of pre-service music teachers’ ratings revealed significant main effects for primary performance area (band, choral, & orchestra) for interpretation/musicianship, F (2, 72) = 4.89, p < .01, partial η2 = .12, and balance/blend, F (2, 72) = 8.68, p < .01, partial η2 = .19. Post hoc tests with the Bonferroni adjustment for multiple comparisons indicated a significant difference between participant groups based on their major area of study (band, choral, & orchestra). Pre-service orchestra teachers assigned more favorable ratings than pre-service choral teachers to the interpretation/musicianship and balance/blend evaluation statements. Pre-service orchestra teachers also gave more favorable balance/blend ratings than pre-service band teachers. A descriptive analysis of the four evaluation statements revealed that pre-service orchestra teachers assigned the most favorable ratings, and pre-service choral teachers gave the least favorable. Although not significantly different from pre-service orchestra teachers, ratings assigned by pre-service band teachers only varied slightly from those given by pre-service choral teachers. Analysis revealed that pre-service orchestra teachers had the largest standard deviations for each evaluation statement, and the largest standard deviation occurred for the other factors evaluation statement. Means and standard deviations for all evaluation statements and primary performance areas are reported in Table 1 (lower means indicate more more favorable ratings).
Pre-Service Music Teachers’ Means and Standard Deviations by Primary Performance Area
Note. Lower means indicate more favorable ratings. Underline indicates a significant difference between primary performance areas. A double underline indicates a significant difference from pre-service orchestra teachers.
Significant main effects for participants’ ratings of interpretation/musicianship, F (1, 72) = 8.47, p < .01, partial η2 = .11, and balance/blend, F (1, 72) = 9.61, p < .01, partial η2 = .12, were found when comparing the two education levels (lower & upper classmen). Ratings given by upper classmen pre-service music teachers to interpretation/musicianship and balance/blend were more favorable than those assigned by lower classmen. For the interpretation/musicianship and balance/blend evaluation statements, an identical rating difference of .33 occurred between lower and upper classmen's evaluations. Although not significant, ratings assigned to dynamics and other factors by upper classmen pre-service music teachers were also more favorable than those given by lower classmen. Means and standard deviations for all evaluation statements and education levels are shown in Table 2.
Pre-Service Music Teachers’ Means and Standard Deviations by Education Level
Note. Lower means indicate more favorable ratings. Underline indicates a significant difference between education levels.
Analysis between the two performance qualities revealed that good performances received more favorable ratings for all evaluation statements when compared to those with a poor quality. Significant main effects were found for interpretation/musicianship, F (1, 72) = 259.10, p < .01, partial η2 = .78; dynamics, F (1, 72) = 152.52, p < .01, partial η2 = .68; balance/blend, F (1, 72) = 186.33, p < .01, partial η2 = .72; and other factors, F (1, 72) = 492.98, p < .01, partial η2 = .87. Means and standard deviations for each performance quality and the four evaluation statements are reported in Table 3. Although ratings assigned to each evaluation statement were significantly different from each other when considering performance quality, it can be seen that the largest variance between the good and poor string orchestra performances occurred for other factors (1.88). Smaller differences were found between the good and poor performances for dynamics (1.06), interpretation/musicianship (1.18), and balance/blend (1.19). Standard deviations for the good performance were smaller on all evaluation statements than those for the poor performance. Analysis revealed that the others factor evaluation statement in the good performance quality condition had the smallest standard deviation (SD = .44).
Pre-Service Music Teachers’ Means and Standard Deviations by Performance Quality
Note. Lower means indicate more favorable ratings. Underline indicates a significant difference between performance qualities.
Discussion
This study was designed to investigate possible effects of primary performance area (band, choral, & orchestra), education level (lower & upper classmen), and performance quality (good & poor) on pre-service music teachers’ evaluations of string orchestra performances. Findings support prior investigations that demonstrated musicians’ abilities to distinguish between performances of good and poor qualities (Byo & Brooks, 1994; Ciorba & Smith, 2009; Geringer & Johnson, 2007; Geringer, Madsen, & Dunnigan, 2001; Johnson & Geringer, 2007; Madsen & Geringer, 1999; Napoles, 2009; Pope, 2012a; Schleff, 1992). The good string orchestra performance received more favorable interpretation/musicianship, dynamics, balance/blend, and other factors ratings than those assigned to the poor performance. The considerable disparity of the two string ensembles may have made it easier for participants to accurately discriminate between the orchestras in the current investigation. When examining evaluations of ensembles that perform at a more comparable performance quality, prior findings revealed that musicians have difficulty distinguishing between the groups. Music majors accurately identified professional and all-state orchestra performances when evaluating technique, but had difficulty when rating musicality (Pope, 2012b). Inconsistent ratings of balance/blend were also found when pre-service and in-service music teachers evaluated average and excellent concert band performances (Doerksen, 1999).
A comparison of past and current results suggest a minimum technique threshold may exist that allows ensembles at different ability levels to produce performances with similar elements of musicality. While all-state orchestras received lower technique ratings than professional orchestras, it appears they had sufficient technical skills to produce performances with high-level musicality characteristics (Pope, 2012b). Future researchers may consider examining the effect of string players’ technical abilities to determine if a minimum level is required to produce performances that receive favorable musicality ratings from adjudicators. Evaluation rubrics used to assess developing string ensembles may need adjustments if findings reveal that young string players need mastery of specific technical skills before adding subjective music elements to performances.
Results from this study support prior investigations that demonstrated adjudicator education level may influence evaluations (Blom & Poole, 2004; Brittin, Sheldon, & Tian, 2002; Fiske, 1977; Geringer, Madsen, & Dunnigan, 2001; Hewitt, 2002, 2005). Analysis of participants’ ratings revealed that upper classmen pre-service music teachers assigned more favorable ratings than lower classmen for all evaluation statements. Caution should be taken when generalizing results from this study since differences were significant only for the interpretation/musicianship and balance/blend evaluation statements. Prior investigators concluded that adjudicators’ accuracy with assigning ratings increased with experience (Ekholm, 1997; Wapnick & Ekholm, 1997), and their findings suggest upper classmen's evaluations in the current study may be more accurate than ratings given by lower classmen. However, Fiske (1975) believed brass players provided more forgiving assessments of solo brass performances than non-brass players due to their understanding of the skills needed to play brass instruments. It is possible upper classmen pre-service music teachers in the current investigation may have gained better insight into the complex skills needed to perform on string instruments during their String Methods course and demonstrated that comprehension by assigning less critical ratings. Lower classmen pre-service music teachers’ less favorable ratings may have resulted from their insufficient knowledge of the specific performance skills needed to play a string instrument.
Analysis of participants’ ratings by primary performance area may also support the theory that increased knowledge of string instruments led to more favorable evaluations. In the current study, pre-service orchestra teachers assigned the most favorable ratings to all evaluation statements and pre-service choral teachers gave the least favorable. Pre-service orchestra teachers more comprehensive knowledge of the skills needed to play a string instrument may have guided them to assign less severe ratings than pre-service band and choral teachers. However, current results differ from prior findings that revealed adjudicators with different performance backgrounds assigned similar ratings to string performances (Geringer et al., 2009; Pope, 2012a; Siddell-Strebel, 2007). The varying backgrounds of participants examined in those studies may help explain contradictory findings. Siddell-Strebel (2007) examined overall evaluations of solo cello performances by elementary school students, adolescents, and adults who did not have formal music training. Undergraduate music majors assigned technique and musicality ratings to all-state and professional orchestra performances in Pope (2012a). Geringer, Allen, MacLeod, and Scott (2009) compared undergraduate music majors and in-service music teachers’ overall evaluations of solo violin performances.
Another possible explanation for the inconsistent results may be the evaluation statements assessed. The current investigation required pre-service music teachers to assign ratings to areas of performance that may be considered relatively subjective (interpretation/musicianship, dynamics, balance/blend, & other factors). Unlike assessments of perhaps more objective elements of music such as intonation, rhythmic precision, and bow direction, adjudicator's personal perceptions of subjective aspects of performance could influence their evaluations. Future researchers may wish to analyze listeners’ perceptions of subjective elements of performance (dynamics, phrasing, balance/blend, vibrato, articulation, & style) to ascertain if contrasts can be numerically quantified. A deeper understanding of listeners’ perceptions of those elements may increase the validity of performance evaluations.
The single presentation order of the string orchestra performances was a limitation of the current investigation and may have affected findings. Ratings assigned to the four evaluation statements may have been influenced by a possible order effect. Future researchers may wish to increase the number of data collection groups and present the recordings in multiple presentation orders in an attempt to control for possible order effects. Due to the limited population and ensembles examined in this study, findings should not be generalized to adjudicators outside of pre-service music teachers and string orchestra performances. Future researchers may wish to extend this study and compare adjudicators’ assessments of band, choral, and string ensembles to reveal if ensemble type and adjudicators’ primary performance areas influence evaluations. To develop a better understanding of the influence of education level on the accuracy of large ensemble evaluations, future researchers may consider comparing the assessments of multiple populations. A comparison of middle school musicians, high school musicians, pre-service, novice, and experienced music teachers may yield different results. Few investigators have examined the effect of string pedagogy courses on musicians’ evaluations of string performances. Participants in the current investigation indicated the completion of String Methods, but those in Geringer et al. (2009) and Pope (2012a) only self-reported their performance backgrounds and did not reveal if they completed string pedagogy coursework. For a better understanding of non-string players’ assessments of orchestral performances, future researchers may consider expanding variables to include secondary instrument performing experience, the completion of university-level string pedagogy courses, student teaching internships, and conference workshops. Understanding the possible influences of non-string players’ learning opportunities on their evaluations of string performances may lead to beneficial changes in teacher education and festival adjudicator training programs.
