Abstract
The Association of Social Work Boards (ASWB) licensing exams play a pivotal role in the social work profession. Administered across the United States and Canada, the ASWB (2023a) states that their exams “test a social worker's competence to practice ethically and safely” (para. 2). Passing an ASWB exam is mandated by nearly all social work state and provincial regulatory bodies, along with completion of several thousand hours of postdegree supervised work experience, as a condition of social work licensure. As these exams play a critical gatekeeping role for social work practitioners, issues of validity, reliability, and fairness of these exams are of paramount importance to test-takers, the social work profession, and the broader public.
Despite the critical role played by the ASWB exams in the profession, scholars have raised concerns about their validity for over a decade (Albright & Thyer, 2010; DeCarlo, 2022; Perron, 2023; Victor et al., 2023). These studies have suggested that the exams do not accurately measure the intended knowledge and skills for social work practice. The accumulated evidence of validity flaws raises serious concerns about whether the ASWB exams serve as a sound measure of a social worker's readiness to practice safely and ethically.
Prior studies have given particular focus to the construct validity of the exams. That is, scholars have raised questions about whether the ASWB exams actually measure the capacity of social workers to practice safely and ethically, which is what the exams are supposed to measure (ASWB, 2023a; Perron, 2023). Albright and Thyer (2010) completed a foundational study in this area nearly 15 years ago. The researchers started with a set of 50 multiple-choice items developed by the ASWB to prepare test-takers for the clinical exam. They then removed the question prompt from each item—leaving only the four provided multiple-choice options—and administered the practice exam to 59 first-year MSW students. Because the question prompts were removed, leaving only the four multiple-choice options, the researchers expected the average score on the practice exam to be 25% based on chance. However, the students achieved an average score of 52%, selecting the correct response twice as often as anticipated. Albright and Thyer suggested that their findings likely indicate a capacity of ASWB test-takers to achieve relatively high scores by capitalizing on cues and patterns present within the multiple-choice answer options, rather than relying on their comprehension of social work principles and best practices. In other words, their study provided evidence of construct-irrelevant variance on the ASWB clinical exam.
Construct-irrelevant variance refers to the presence of extraneous factors that can influence test results, which are unrelated to the construct being measured (Haladyna & Downing, 2004). This variance can pose a significant threat to the validity of test scores, as it can lead to inaccurate inferences about the abilities or traits of the individuals being assessed. When construct-irrelevant variance is present, it undermines the ability of a test to accurately capture the targeted construct—in this case, the capacity of social workers to practice safely and ethically—resulting in misleading or biased outcomes. In order to ensure the validity of test results, test developers should minimize construct-irrelevant variance by carefully designing test items, providing appropriate accommodations, and considering potential sources of bias during test development and administration.
Current Study
Given the passage of time and potential exam changes, this replication study investigates whether the construct-irrelevant variance identified by Albright and Thyer (2010) persists in the current version of the ASWB clinical exam. By reassessing this issue with newly developed items, we aim to determine whether serious flaws related to construct validity remain, which would potentially render the exam an ineffective and inaccurate tool for evaluating the competencies of aspiring social work professionals. We use the same approach in this study, deploying ChatGPT, a large language model (LLM), to assess current ASWB-developed materials and determine whether the structure of multiple-choice options continues to serve as a source of construct-irrelevant variance on the clinical exam. ChatGPT has previously been used to evaluate licensing exams in both social work (Victor et al., 2023) and other professional disciplines including medicine (Gilsen et al., 2023) and engineering (Pursnani et al., 2023).
Method
Simulated ASWB Clinical Exam
In the current study, we used a set of 50 ASWB-developed exam items provided to educators to help to prepare students for the actual ASWB clinical licensing exam (ASWB, 2023b). The items we used for the study appeared on previous versions of the ASWB clinical exam, and are currently distributed by ASWB as exam preparation materials. We therefore assume that these items mirror those actively in use on the ASWB clinical exam, and are thus a strong proxy for current exam items.
Design and Procedures
We utilized a comparative procedure to assess the performance of an LLM in completing two different formats of the clinical exam. The first format was traditional, where the model was presented with both the question and the four multiple-choice responses for each exam item. The second format was a modified version based on Albright and Thyer's (2010) design. In this version, only the four multiple-choice options were provided, with the question itself being omitted. This procedure helps us understand the extent to which the question context is important in selecting a correct answer and to determine if the wording of the multiple-choice options continues to serve as a source of construct-irrelevant variance.
For this study, we used ChatGPT-4 (May 24 version), one of the most powerful and widely used LLMs. Utilizing ChatGPT to perform this replication study presents some advantages and weaknesses compared to administering these items to first-year MSW students. A primary advantage is that ChatGPT is not influenced by factors such as fatigue, stress, or personal biases that might potentially affect a human test-taker's performance. This allows for a more controlled and consistent test-taking environment, contributing to the reliability of the study's results. Secondly, ChatGPT does not possess any specific instruction in social work, and can only draw from the texts on which it was trained. Therefore, its performance on the ASWB licensing exam can be attributed to its ability to detect patterns within the answer choices. We consider the weaknesses of using ChatGPT below in the limitations section.
We used ChatGPT to complete both the traditional and modified version of the exam. For the traditional version with the question included we used the following prompt to ask ChatGPT to provide an answer to each of the 50 clinical exam items in our data set: Assume you are a first year MSW student asked to take the ASWB licensing exam. I will give you a set of exam items with both the question prompt and the multiple-choice options. Please provide me with the letter of the response you believe to be the correct answer for each exam item.
For the modified version we used the following prompt designed to mimic the original study by Albright and Thyer (2010) and to collect information on the decision-making strategies of the model when the question prompt was not available: Assume you are a first year MSW student asked to take a unique version of the ASWB licensing exam. You will be given multiple-choice options without the actual questions. I will give you each set of multiple-choice options. It is difficult to identify the best option without the context. Please do not provide an analysis of the options. Instead, do your best to figure out what the best answer could be. After you figure out the best response, provide me with a description of your decision-making process. Your output should be in the format:
{Answer}, {Explanation}
ChatGPT is a probabilistic model, meaning that it generates responses based on probabilities assigned to different sequences of words, and in this case, probabilities are assigned to the likelihood that a multiple-choice option is the correct answer. When the model is confident in a given answer it is likely to select the same choice during repeat testing. However, when the model assigns probabilities that are close to one another, it may change its answer when prompted a second or third time. To ensure the integrity of the results and assess for variation in performance, we completed three rounds of testing for the modified version of the exam. For each round, we started a new chat and began with the prompt above.
Analysis
We assessed performance with a simple count of the number of questions that each model answered correctly. To score the actual clinical exam, ASWB uses a pass-point approach. Under this approach, the number of correct answers needed to pass the exam varies between 90 and 107 out of 150 scored items depending on the version of the clinical exam being administered (ASWB 2023c). Thus, test takers must correctly answer between 60% and 71.3% of items to pass depending on the version's pass point. We first determined the model's performance across exam items with the question prompt included to serve as a point of reference. We then calculated the models’ performance (i.e., accuracy rate) across three rounds of testing when the question prompts were removed in order to estimate the extent to which current responses might serve as a source of construct-irrelevant variance.
Results
To establish a point of reference for considering the level of construct-irrelevant variance, we first determined the accuracy of ChatGPT when taking the traditional version of the clinical exam with both the question and multiple-choice options available for each item. Under those conditions, ChatGPT correctly answered 90% of the items on the clinical exam. We then conducted three rounds of testing on the modified version of the clinical exam, asking ChatGPT to select the most likely multiple-choice answer without the question available. The model accuracy was 70%, 76%, and 74% across these three rounds. Fleiss’ kappa across the three rounds was 0.81 indicating almost perfect agreement (Landis & Koch, 1977).
Discussion and Applications to Practice
The findings of this study—which replicate those of Albright and Thyer (2010)—indicate that ChatGPT, an artificial intelligence LLM, was able to perform far better than chance on the ASWB clinical licensing exam when the questions were hidden and only the four multiple-choice options were presented. Instead, ChatGPT achieved an average accuracy rate of 73.3% across three rounds of testing. In two of these instances, model accuracy was high enough that ChatGPT is likely to have passed the clinical exam without the questions provided.
These results raise significant concerns regarding the construct validity of the ASWB clinical exam, as it suggests that the exam might not be measuring the intended construct of social workers’ competence to practice safely and ethically. Instead, it appears that relatively high scores can be achieved on the clinical exam by relying on cues and patterns present within the answer choices themselves, a strong indicator of construct-irrelevant variance. This observation provides further support that use of the current ASWB clinical exam to evaluate the competence of social work professionals is not likely to be appropriate for its intended use. In light of these findings, the developers and administrators of the ASWB licensing exams must reassess the exam's structure, with the aim of eliminating potential biases and improving its construct validity. This may involve the use of alternative question formats, the implementation of more rigorous item-writing guidelines, or—our preference—the development of novel assessment methods that better capture the knowledge and skills required for safe and ethical social work practice.
Fortunately for the profession and the broader public, clinical social work licensure requirements already include extensive supervision by a licensed practitioner. In terms of assessing the construct of the ability to practice safely and ethically, supervisors are likely better equipped than a multiple-choice exam to determine the true capabilities of those seeking licensure. That is, observing performance in context is a more valid approach to assessing competence to practice safely and ethically (Haladyna & Downing, 2004). Supervisors are able to go beyond the evaluation of theoretical knowledge about professional ethics and safe practice, as assessed in licensing exams. Supervisors can provide real-time, situational guidance on safety and ethical dilemmas that arise in the field, fostering a nuanced, practical understanding of principles and practice guidelines that can be more contextually relevant and adaptive. This in-context approach can nurture the ability to apply ethical frameworks to complex, real-world situations in a manner that the abstract and/or information-limited questions on the licensing exam might not fully capture.
At least one jurisdiction has adopted an alternative pathway to clinical licensure that places more reliance on the clinical supervision model currently mandated by social work regulatory bodies. The Illinois Legislature recently passed House Bill 2365 which provides a pathway to clinical licensure that does not require passage of the ASWB clinical exam (Illinois General Assembly, 2023). Applicants are required to take the clinical exam once, and those who do not pass can opt for an additional 3,000 hours of supervised practice to obtain their license.
If testing is to remain a core feature of licensure, then we strongly advocate for increased transparency and independent evaluation of the ASWB exams to ensure conformity to the Standards for Educational and Psychological Testing. This publication, jointly produced by the American Educational Research Association, American Psychological Association, and the National Council of Measurement in Education (2014), details the specific criteria for developing and evaluating testing practices and for assessing the validity of score interpretations. The publication includes a special section that focuses specifically on professional credentialing. In particular, Cluster 1 Standards (11.1, 11.2, and 11.3) and Cluster 3 Standards (11.13–11.16) provide explicit guidelines concerning validity evidence for professional credentialing. Social work is not alone in this need for independent evaluation of its licensing exam, as the need for increased transparency and independent validation of exams has been called for in other mental health professions such as clinical psychology, counseling, and marriage and family therapy (Caldwell & Rousmaniere, 2022).
The recent announcement by the Association of Social Work Boards (ASWB, 2022a) regarding their decision to reduce the number of multiple-choice options on exams from four to three exemplifies the pressing need for transparency and independent evaluation in our field. The organization asserts that this exam modification stems from “psychometric expertise that confirms the validity” (para. 2) of this reduction in multiple-choice options. However, a significant issue arises as ASWB provides no publicly available evidence to back its claim.
The responsibility to accept ASWB's assurance about the validity of exam adjustments falls on test-takers, educators, and policymakers. However, a clear conflict of interest makes this faith-based acceptance worrisome. The conflict arises because the exam developer, ASWB, also assumes complete control over establishing and endorsing the exam's validity. This arrangement carries a potential bias, given ASWB's inherent interest in promoting the effectiveness and validity of its exam which serves as the organization's primary revenue and profit generator. In 2021, the most recent year for which revenue data are available, ASWB reported profits of over US$7 million, up approximately US$4 million from the year before (ASWB, 2021). Such a conflict of interest introduces considerable bias into the evaluation process, ultimately undermining the exam's validity. This conflict of interest highlights the importance and need for independent third-party verification to address the conflicts and bolster confidence in the validity of any modifications made to the test format.
While using ChatGPT for this study had advantages, it also presented limitations. First, as an AI language model, ChatGPT's decision-making might not mirror that of human test-takers. The model selects answers based on patterns and statistical associations that likely differ in some ways from how humans would arrive at an answer. This could mean that the selections the model makes when faced with incomplete information may not align with a human test-taker's selections under the same conditions. Secondly, ChatGPT has not received explicit training in social work knowledge and skills which could also be considered a disadvantage in generalizing these findings to authorized test-takers. Workers who have received advanced social work training and are engaged in clinical practice might make different decisions under these circumstances when drawing on their content knowledge or work experience.
In their original article documenting construct-irrelevant variance on the ASWB clinical exam, Albright and Thyer (2010) ask in the subtitle: quis custodiet ipsos custodes? In English, “who is watching the watchers?” Fifteen years later, the answer does not appear encouraging. Major validity and transparency concerns remain unresolved (DeCarlo, 2022; Victor et al., 2023), and the use of the exams in the licensure process continues to exclude workers from the profession based on race, age, and first language (ASWB, 2022b). Given that the ASWB has yet to correct the long-standing validity challenge replicated in this study—in addition to other identified flaws—we encourage state legislatures and regulatory bodies to temporarily suspend the use of the ASWB clinical exam as a requirement for licensure until an acceptable and independently validated alternative is established.
Footnotes
Acknowledgments
The authors would like to thank Dr. Matthew DeCarlo at Saint Joseph's University for his helpful review and feedback on this manuscript. We acknowledge the use of ChatGPT-4 (May 24 version) in the writing stages of this article, both the original draft and reviewing/editing, and in formal analysis. ChatGPT-4 was used for generating text, manuscript editing, and analyzing text data which was central to this study. We have carefully reviewed the AI-generated content to ensure its accuracy and validity, and we assume full responsibility for the content presented in this article.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
