Abstract
In Taiwan, both the academic and practitioner have noticed the disputation of high-rank civil service recruitment is getting worse due to the oral exam method. In this article, there are two research questions that need to be answered. One of them is, “How does the oral exam influence the test scores and to what degree?” The other is, “How can the oral exam be improved?” To answer these questions, the research methods of content analysis, Delphi technique, and in-depth interview were adopted. To provide persuasive evidence to disclose the serious problems of oral examination, our research chose four kinds of civil service exams as research objectives. The results of our research not only demonstrate measures for the reform of oral exam assessment but also reveal characteristics of oral exam assessment in Taiwan. Through these results, we get a better understanding of the stakeholders’ perceptions about oral examinations. Besides, it is helpful for the emergence of reinventing strategies.
Behind the Scene
Every year in Taiwan, about 60 national civil service exams are held; all examinations have both traditional written exams as well as oral exams that make up for about a third of the examination. According to oral examination guidelines, the oral examination is divided into individual oral examination, collective oral examination, and group discussion. 1 The results of the oral examination account for about 10% to 20% of the total examination grade, causing for some controversy about the oral examination’s influence on admittance results. Apart from some examinees’ complaints and even lawsuits on the matter, people also questioned about its procedure, fairness issue, transparency, professionalism, and other personnel administration considerations. This has made it a hot topic of research in Taiwan.
After reviewing documents dealing with the method of Taiwan’s civil service exams, we have put together a list of the different complaints concerning the oral examinations grading system: “The oral examination questions are not predicated on job analysis,” “Similar oral exams are the responsibility of several different groups of oral exam commissioners,” “The time for the oral exam was too short,” “There lacks a constant standard for grading oral exams,” “Oral exam commissioners lack professional training of interview skills,” “The possibility of unfair treatment towards those taking the exams,” “The oral exam commissioners may influence each other,” “There is not a well established oral exam database of questions,” and “The group discussion part of the examination lacks professional design” (Huang, 2007; Wang, 2004; Wu, 2000, pp. 50-51; Wu, 2007, pp. 16-18). It can be seen from the analysis above that the oral examination of Taiwan’s civil service exam process is lacking in structure.
In order to clarify the points of controversy mentioned above, we need to assess the impact of the oral examination grades on examinee’s admission. At the same time, we explain some of the more common reforms being made in the area of grading oral examinations. And finally, we identify the oral examination traits that exist in Taiwan. In other words, we attempt to make a response to the need of improving oral examination’s reliability and validity by performing a structural analysis of related civil service examinations.
Theories and Literatures Reviewed
The oral examination is either structural or nonstructural, which has a large impact on exams’ reliability, validity, and its results (Campion, Palmer, & Campion, 1997). For this reason, many researchers of oral examinations have started to look into the difference in influence of structured oral exams and nonstructured oral exams on reliability and validity (Campion, Campion, & Hudson, 1994; Campion et al., 1997; Motowidlo et al., 1992; Wu, 2007). After an analysis of different scholar’s definitions of highly structured oral examinations, we can form a common definition: “an oral exam in which the questions are based on job analysis, and the content and procedures of the exam are standardized” (Campion et al., 1997; Huffcutt & Arthur, 1994; Motowidlo et al., 1992).
The biggest difference between structured oral exams and nonstructured oral exams is that the question designs of structured oral exams are based on job analysis, and the content and procedure of the exam are highly institutionalized and standardized. Therefore, structured oral examinations lessen the impact that the subjective opinions and personal background of the oral exam commissioners might have on the admittance results of oral examinations (Graves, 1993), and in doing so, improve the fairness of the oral examination.
Structured oral examinations are being promoted among the academic and practical circles, and are highly regarded as a more reliable and valid method of oral examination. How to progress toward structured oral examinations has become a very important topic in the practice of oral examination. How is the structuralization of oral examinations defined? The research of Campion et al. (1997) defined the structuralization of oral examinations as “any enhancement of the interview that is intended to increase psychometric properties by increasing standardization or otherwise assisting the interviewer in determining what questions to ask or how to evaluate responses” (p. 656). Huffcutt and Arthur (1994) defined the structuralization of oral examinations as “lowering the level of variation among examinees during the interview process, and explaining it as the level of decision making by the interviewer” (p. 186). Motowidlo et al. (1992) defined the structuralization of oral examinations as “the amount of freedom the interviewer has in making decisions.”
Through the definitions given above, we can see that the structuralization of oral examinations refers to the increase of standardization in the examination process and behavioral model of the examination officer, thus lowering the amount of variance occurring during the examination process. Basically, the higher the level of standardization, and the lower the possibility of variance, the higher the structuralization of the oral examination is, and vice versa.
What are the traits of structuralized oral examinations? After categorizing the definition of structured oral exams by the following researchers (Blackman, 2002; Campion et al., 1994; Campion et al., 1997; Campion, Pursell, & Brown, 1998; Chapman & Zweig, 2005; Conway, Jako, & Goodman, 1995; Dipboye & Gaugler, 1993; Huffcutt & Arthur, 1994; Hysong & Dipboye, 1998; Motowidlo et al., 1992; Pulakos & Schmitt, 1995), we discovered that there are a total of eight characteristics: (a) standardization of questions, (b) standardization of response, (c) standardization of rating scales, (d) base questions on a job analysis, (e) multiple dimensions of rating, (f) voices from candidates, (g) establishment of relationships, and (h) taking detailed notes (see Table 1). The categorization above shows what most scholars identify to be the four essential traits of oral examination assessment: “standardization of questions, standardization of response, standardization of rating scales, and base questions on a job analysis.” This implies that the essence of carrying out structured oral examinations is the content of the oral exam (What questions are asked?), the grading scale used (What is the standard for awarding points/grades?), the procedure of the oral examination (How is the examination planned?), and so on.
Critical Features of Structured Interview.
Through analysis of the results of the research mentioned above, we find that structuralization can improve the reliability and validity of an examination. Campion et al. (1997, p. 657) proposed 15 ways that can increase the structuralization of oral examinations 2 ; furthermore, research has shown that most oral exam structuralization is helpful for the increase of reliability and validity of the exam and can even be seen in the reaction of the two parties involved in the oral exam. Also, the previously mentioned analysis serves as a reminder that on a practical level, there can also be an occurrence of a negative-impact relationship, for example, “limit prompting has negative effect on candidate reaction.”
Research Design
Research Framework
After a thorough analysis of the main theories and current information regarding oral exam assessment, we found that there are seven essential points that influence the oral exam assessment method:
The object of oral exam assessment: There is much discussion and differing of opinions concerning the object of oral exam assessment on both a theoretical and practical level; for example, “Should research favor the measure of personality traits or of professional knowledge to distinguish the different functions of oral and written examinations?”
The assessment method of oral exam: Apart from the object of oral exam assessment, both theoretical and practical research emphasize the importance of the discussion of the assessment method of the oral exam. Related arguments can be separated into two groups—structured oral exams and nonstructured oral exams. The former emphasizes that oral exams should be based on job analysis and that the content and procedure should be standardized, whereas the latter, from test problem design to content and procedure, has no complete standardization to speak of.
The composition of the oral exam commission: An analysis of documents shows that the selection of oral exam commissioners, the background of oral exam commissioners, the training of the oral exam commissioners, and other matters have a deep impact on the effectiveness of oral exam assessment. So, if we want to truly understand the improvement strategies of oral examination’s methods and procedure, we must first analyze the formation of the oral exam commission.
The standardization of oral exam procedure: Time allotment, record of dialogue, grading tools, oral exam preparation meetings, the standardization of the process in which the oral examination committee asks questions, the standardization of questions, the standardization of answers (including grading standards, calculating of scores, and grading statements), and so on.
Types of oral exams: According to “The Rules of Oral Exam” approved by the Examination Yuan of Taiwan government, current types of oral exams can be separated into three categories: individual oral exams, collective oral exams, and group discussions. The emphasis of assessment and the procedure of each type are different. Individual oral exams emphasize the candidate’s demeanor, language use, and knowledge. The collective oral exam increases the assessment criteria of responsiveness. Group discussions emphasize one’s amiability, receptiveness, decisiveness, and so on. 1
Grading categories and the ratio of grade allotment: Taiwanese civil service tests are all performed according to a set of rules; the general principle of the oral exams can be found within “The Rules of Oral Exam,” but individual tests in accordance with these rules establish their own exam guidelines, whose content clearly establishes grading categories and grading scales. According to our analysis of examinees’ admittance results, there is a definite proportion of influence caused by the oral exam score. Therefore, this should be taken into account when considering how to improve the method of oral examinations.
Other regulations: Analysis of documentation shows that the content includes the right of speech of the examinees, confidentiality, avoidance of interests, and legal regulations (sex equality laws, disability laws, employment service laws, etc.).
This article was based on the above analysis, to establish an analytical framework for the follow-up “video record analysis” and “Delphi Survey” (refer to Figure 1). The purpose of applying Delphi here is to verify, revise, and propose strategies with consensus to improve the oral exams’ reliability and validity for Taiwan civil service system.

Analytical framework.
Research Methods
In this article, we perform a documentation review, content analysis, and Delphi survey, in that order. First, we used four steps to choose four classes of exams from a pool of 22 different applied oral exams each using a different assessment method. 3 Then, we performed content analysis of the civil service oral examination assessment video recordings for the year 2007 to increase our understanding of the current situation of the Taiwanese oral exam assessment method, to form a research structure and Delphi survey. The object of the Delphi technique includes policy makers, oral examination commissioners, and examinees; those selected participated in an in-depth interview 4 to draw up a conclusion of the discoveries and suggestion of the overall research (Table 2).
Delphi Technique Assessment Dimensions and Content.
For the survey answering system, we used the Likert-type scale with response options “strongly disagree,” “disagree,” “no opinion,” “agree,” and “strongly agree.” The grading scale ranged from 1 to 5 points—The higher the score, the more that respondent agrees with the relevance of the index. At the same time, respondents were given freedom to give their opinion toward every index; if the respondents chose “strongly disagree” or “disagree” as their response to a certain question, they had to write their opinion of how it should be revised. If the respondents felt that there were any composition questions in the previous questionnaire that could be added or deleted, they were given a place in the “revision opinion column” and “overall opinion column” to voice these opinions.
Delphi survey was applied in this study for problem analysis and strategies making. As for the research design of Delphi survey, the method of purposive sampling was used to select the subject of the 75 respondents 5 in the first run survey, which were divided into three groups of stakeholders: 62 entry-level civil servants who gain admittance by passing oral exams, 4 senior public officials of department heads, and 9 national exam oral exam commissioners. Under given research design, all respondents were divided into two groups (examinees and nonexaminees) for further analysis and discussion. We totally conducted two rounds of Delphi survey, which was helpful in collecting suggestions for improvements made about the oral exams. At the same time, we performed an analysis of the video records to find the factors that influenced reliability and validity, and also to compare with the results of the Delphi survey. Finally, we interviewed the professionals/scholars as well as the government policy makers or executive officers to get their opinion and analysis on oral exam topics that were somewhat controversial.
We used “standard deviation” to measure the difference in opinion among the different respondents. When the standard deviation is below 0.6, it could be seen as “high level of consensus” among respondents; when the standard deviation is between 0.6 and 1, it means a “medium level of consensus”; and if the standard deviation is above 1, it presents a “low level of consensus.” Also the mean was used to determine the respondent’s level of agreement with a certain index; the higher the mean, the higher the level of agreement.
Research Analysis
For this research, we selected four types of Taiwanese civil service examinations, 6 and performed an analysis of their oral exam questions and the assessment methods of the oral exam commissioners. We found that under current guidelines, the oral exam scores did not make up a large proportion for the overall admittance grades (only about 10%-20%) but was still influential for the exam. We calculated the admittance results of the four types of exams during 2005 to 2007 to better understand “after the adding of the oral exam scores, the degree to which the overall rank of the examinees was influenced.” 7 As a result, we found that in 2005, for the four types of exams, after adding the oral exam scores, the percentage of movement of the examinees was about 35.71% to 72.24%. In 2006, the results were between 25.4% and 72.08%, and in 2007, the results were between 38.39% and 72.22%.
According to the research design mentioned above, we then performed “the content analysis of oral exam video recordings” and “Delphi method survey” to identify the influential factors for the oral exams. Finally, a strategy for improvement was formed. Below is an in-depth explanation of this research.
Content Analysis of Oral Exam Video Recording
The research team surveyed the video recordings of the four exams that were provided by the government authority. We applied the analytical framework established after the literature review, as fundamental criteria for video recording analysis, to explain which factors may have influenced the structure of the oral exam and the level of influence that these factors had.
It can be seen from Table 3 that the factors that may have influenced the assessment of the four types of oral exams were “follow-up questions,” “unrelated questions to professional or job-related knowledge,” “use of ancillary information,” “did not allow everyone to complete the supplementary answers,” “interaction between examinees and oral exam commissioners,” “time management problems of the oral exam commissioners,” “difference in the amount of questions asked among the different type of exams,” “questions that were specified toward one individual examinee,” and so on. After further analysis of research results of individual cases, it can be seen that the factors that generally influence the structure, standardization, and fairness of oral exams are mainly “allowing examinees to give supplement answers,” while deriving the factor of “not allowing every person to complete the supplement answers.”
Factors That Influenced the Oral Exam Structure and Frequency of Occurrence.
The oral examiner did not take part in the group discussion procedure, but rather stood to one side assessing the performance of the examinees.
During group discussion, the oral examiner does not participate, but stands to the side and assesses the performance of the examinees. “Like-me effect” refers to the attitudes between the examinee and the oral exam commissioner when the two share similar background, resulting in higher scores for the examinee. For example through content analysis of the oral exam, it can be seen that sometimes when the oral exam commissioner is aware that the examinee is a graduate from the same school, or is a student of their school, his or her attitude and the content of dialogue will be much more warm and easygoing, while being more strict with other examinees.
“The interaction between examinee and oral examiners” refers to the situation where there are unnecessary interactions between the interviewer and examinee during the oral exam.
Comparing with the overall analysis listed above, different oral exams have their own problems. In terms of “collective oral examination,” problems such as “During the test the oral exam commissioners asked follow-up questions; The oral exam commissioners provided question asking, but the examinees did not ask questions; Used supplemental information; The oral exam commissioners had trouble with time management; Questions were asked which were unrelated to professional knowledge; Examinees were allowed to give supplemental answers; The interaction between examinees and exam commissioners,” and so on usually occur. Besides, the result presents that “individual oral examination” often encounters challenges such as “The exam officer allowed opportunities for asking questions, but the examinee did not ask questions”; “During the exam the oral exam officer asked follow-up questions”; “Questions were asked that were directed toward one individual examinee”; “Allowed examinees to give supplemental answers”; “Use of supplemental information”; and so on.
Content analysis shows that the 10 factors shown above were mainly responsible for the influence of the oral exam structure, reliability, and validity. When deciding on how to design the Delphi survey, four important points were considered:
1. The time management ability of the oral examiners
Content analysis reveals that the time management ability of the oral exam commissioners was not up to par, and that this problem further derived additional problems that influenced the structure, reliability, and validity of the oral exam, including “The oral exam commissioner provided unequal opportunities to ask questions” and “The examinee was given an unequal opportunity to give supplemental answers.” These situations usually were a result of poor management of time by the oral exam commissioner. Moreover, it might cause the oral exam to be performed roughly, and the structure of oral exam could be broke. Therefore, the time management training for the oral exam commissioners is a topic that needs to be discussed.
2. The use of ancillary tools
Questions that were unrelated to jobs, interaction between examinees and oral examiners, and questions asked that were geared toward individual examinees were examples of situations resulting from the use of supplementary tools by the oral examiners, which influenced the reliability and validity of the oral examination. It is possible that access to supplementary tools allowed oral examiners to know about the educational and professional experience of the examinees, to form prejudices about the candidates, and to spend time asking unrelated questions. Therefore, it is also important to discuss the necessity of ancillary tools.
3. The standardization of the rules in asking questions
Content analysis revealed that the oral exam questions were asked quite disorderly. Some examinees were asked different questions, whereas others were asked the same questions depending on which exams they took and who they encountered. In the end, this situation had disturbed the reliability and validity of the oral exam. Therefore, interviews should be used to further discuss the standardization of the subject matter of oral exam questions.
4. Types of oral exam questions
Some oral exam questions are based on professional knowledge, others are based on practical experience, while yet others are totally unrelated to professional knowledge or job experience. However, the purpose of the oral exam should be to explore the suitability between each examinee’s character traits and job, as professional knowledge is already tested in the written part of the exam. It can be said that professional knowledge testing in the oral exam might cause a lot of overlapping, taking away the supplement effect that the oral exam should have; this point is also worthy of further discussion.
The Analysis of First-Round Delphi Survey
The design of first-round Delphi survey was based on the literature analysis. We established three-dimension assessment indexes as a frame of questionnaire. The survey questionnaires were sent out through mail on November 30, 2008, and were received beginning December 7, with the first round of surveys received by December 12. There were a total of 75 questionnaires sent out during the first round, 50 were returned, for a return rate of 66.7%; the results of the surveys can be seen in Table 4.
The Return Rate of the First-Round Delphi Survey.
We used “standard deviation” to measure the difference in opinion among the different respondents. Also the mean was used to determine the respondent’s level of agreement with a certain index; the higher the mean, the higher the level of agreement. Of the 40 questions asked in the first round of survey, there were 6 questions that reached high level of consensus (SD below 0.6), 34 questions that were at the level of medium level of consensus (SD between 0.6 and 1), and there were no questions with a low level of consensus. This result could be a primary justification of our previous literature analysis and for the indexes of Delphi survey.
The Analysis of Second-Round Delphi Survey
The purpose of second-round survey was to justify further the suitability between respondents’ opinions and the indexes that were established through literature analysis. The results of first-round survey were sent to the same experts for reference when they made decision again. Also, the survey is a second opportunity for those stakeholders to answer and express their different opinions. This process allows the respondents to understand the similarities and differences of each other’s opinions. The second-round survey questionnaires were sent out through email on December 17, 2008; beginning on December 24, the survey questionnaires were gradually received. The receiving of surveys ended on January 8, 2009. A total of 50 questionnaires were sent out, 37 were received back, with a return rate of 74%. The return rate of the survey by different types of respondents can be seen in Table 5.
The Return Rate of the Second-Round Delphi Survey.
Initial analysis of the second-round survey shows that, other than a few exceptions, most of the indexes showed that the standard deviation dropped, while the mean rose. Of the 40 questions, 14 reached the level of high consensus (SD < 0.6), several more than the 6 questions in Round 1, showing that the respondents’ opinions became more consistent, and the level of agreement concerning the main dimension, subdimension, and indexes increased.
As for the three main dimensions—“the standardization of oral exam procedure (SD = 0.55),” “the reliability and validity of oral exam methods (SD = 0.62),” and “operational regulations in practice (SD = 0.71)”—respondents’ opinions of “standardization of oral exam procedure” changed from the “medium level of consensus” to a “high level of consensus” in Round 2; the other two dimensions remained at a medium level of consensus. However, all the means displayed a rising pattern, showing that the level of agreement among respondents regarding these three dimensions increased.
Comprehensive Analysis of Delphi Survey
After two rounds of Delphi survey, we found out the subdimension’s order of agreement among respondents were “types of oral exams (SD = 0.51),” “operational issues (SD = 0.52),” “assessment criteria of oral exam (SD = 0.53),” “oral exam procedure (SD = 0.56),” “standardization of oral exam (SD = 0.56),” “regulatory issues (SD = 0.57),” “suitability of oral exam methods (SD = 0.60),” and “formation of Oral Exam Commission (SD = 0.63).” Although the SDs are not as low as the three main dimensions, it also reveals that both examinees and nonexaminees feel that the reliability and validity of the oral exam are very important, and they both feel that “the standardization of oral exam procedure” is the essential dimension for a better oral exam system. Therefore, “the composition of the oral exam commissioners” and “the standardization of the procedure” are becoming very important factors in reaching the goals of this research.
Under the main dimension of “the standardization of oral exam procedure,” there is a total of seven indexes. The top three in level of agreement among respondents were, “There should be a regulation for the training of oral exam commissioners (SD = 0.55),” “There should be an oral exam preparation meeting held before the interview (SD = 0.62),” and “The process for selecting the oral exam commissioners should be standardized (SD = 0.63).” These results revealed the importance of the preparation matter before interview has been emphasized by respondents.
Under the main dimension of “the reliability and validity of oral exam methods,” there was a total of 11 indexes. Of these 11, the top 3 with the highest level of agreement among respondents were, “The assessment criteria of individual oral exams should emphasize demeanor, language use, and knowledge (SD = 0.51),” “The content of oral exams should include the assessment of professional ability (SD = 0.55),” and “The content of oral exams should include the character traits (SD = 0.62).” These results displayed a paradox between “The Rules of Oral Examination” and respondents’ subjectivity. Meanwhile, we noticed that there was a great difference among respondents on the indexes of “the assessment criteria of group discussions (SD = 0.82)” and “the proportion of oral exam for the overall grade allotment (SD = 0.84).”
The main dimension of “the operational regulations in practice” had a total of 11 indexes. Of these 11, the top 3 in terms of level of agreement among respondents were, “During the oral exam, the avoidance of interest between examinees and interviewers should be carried out (SD = 0.49),” “The guidelines of oral exam should be prepared and provided to all interviewers before the exam (SD = 0.58),” and “The Judicial Personnel Level 3 Exam should adapt both the individual and collective oral exams (SD = 0.59).” These results show that from an operational regulations aspect, respondents feel that the avoidance of interest relationship is a critical factor for a better oral exam system. Besides, the opinions of this portion of the survey are highly related with the fairness of the oral exam.
Through the results of the second-round Delphi survey, the importance of three main dimensions established in this research have been accepted by the respondents. Moreover, the opinions among the respondents have begun to converge. This shows that the respondents feel that reliability and validity are the most important factors in achieving an objective oral examination.
Research Findings
After comparing the results of the content analysis of the video recordings and the Delphi survey, it seems that the analysis mentioned above is just another verification for the claims of standardization for oral exams. The most interesting findings while using the Taiwanese civil service exams as the research object are the discovery of two characteristics of oral exam assessment and one contradictory opinion among respondents, which are described below:
The examinees tend to be silent when they are allowed to voice: According to the Delphi survey results, respondents felt unfairness of the right of speech among different examinees. This result encourages us to pay attention to the video recordings analysis. We expect to gain more understanding of the fairness issue of the right of speech during the oral exam. After reviewing the video recordings, we found that there were many occurrences of oral examiner giving the examinees chances to ask questions or perform speeches. But the chances were not equally given to each examinee.
The examinees do not place a strong emphasis on the avoidance of interests: Through examination of the video recordings, we found that there were some examinees who purposefully emphasized their school, work, or teacher in an attempt to stress their background to get a higher score. This situation is not in accordance with the principle of avoidance of interests. The results of the Delphi survey show that the index, “During the oral exam, the avoidance of interest between examinees and interviewers should be carried out,” was identified with high level of consensus by the respondents. We also found that respondent group made up of public officials and experts/scholars had a higher level of agreement in this area than the examinees. But it does not seem consistent with the results of the video recordings. Therefore, during oral exam preparation meetings and training, the problem of the avoidance of interests should be paid special attention.
There exists a contradiction as to whether question asking should be standardized: According to literature, “the standardization of oral exam commissioner question asking procedure” is important to increase the standardization of the oral exam. But through the content analysis of the video recordings, we found that the questions asked did not meet the requirements of standardization. Also, the Delphi survey results found that some respondents felt that “The procedure of asking questions and content of questions asked in an oral exam should not be standardized in avoidance of lacking flexibility.” In light of the theoretical and practical contradiction mentioned above, this would be a good topic to do further research on in the future.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
