Abstract
This study examined the Word Choice Test’s (WCT) utility as a performance validity test in a mixed clinical sample of veterans referred for neuropsychological evaluation. Participants completed Green’s Word Memory Test (WMT), WCT, and Test of Memory Malingering (TOMM) Trial 1. Using the WMT as the criterion for valid performance, logistic regressions examined the WCT and TOMM’s classification accuracy for those with and without cognitive impairment (CI). Receiver operating characteristic curves were used to establish cut scores which maximized the sensitivity/specificity of each measure. In those without CI, both tests showed good classification accuracy (86.7% and 85.0%, respectively). Among those with CI, the TOMM retained good classification accuracy (82.3%), while the WCT’s decreased considerably (69.4%). Optimal WCT cut scores differed based on impairment status, with a higher sensitivity/specificity trade-off among those with CI. Successful performance on the WCT appeared to rely more heavily on cognitive processes unrelated to performance validity.
Over the past 30 years, recognition of the need for assessment of performance validity in the context of neuropsychological evaluations has grown, and research in this area has burgeoned accordingly. There are now a multitude of standardized and embedded performance validity tests (PVTs) at neuropsychologists’ disposal, with varying amounts of empirical support. This proliferation of PVTs has been essential to meet professional calls and standards of practice for the assessment of performance validity via multiple methods in every evaluation (Bush et al., 2005; Heilbronner, Sweet, Morgan, Larrabee, & Millis, 2009).
To reduce both time and health care expenditure burden in neuropsychological evaluations, it may be preferable to administer PVTs that are brief in duration and demonstrate specificity and sensitivity similar to their longer counterparts. The Advanced Clinical Solutions (ACS) Word Choice Test (WCT; Pearson, 2009) was recently developed to meet such a need. The WCT is a forced-choice measure of performance validity, as assessed by recognition memory for a 50-item word list presented visually and orally. The initial validation was conducted in a clinical sample of 371 patients with diverse neurological, psychiatric, and developmental disorders. Base rates representing the cumulative percentage of the sample scoring at or below various numbers of correct responses were reported for the sample as a whole, as well as for each diagnostic subsample. However, there was no external criterion of performance validity against which WCT classification was compared. Thus, the ability to draw conclusions about the adequacy of task engagement was limited. Moreover, while individuals with moderate to severe traumatic brain injuries (TBIs), temporal lobectomies, learning disorders, and attention-deficit/hyperactivity disorder were included in the original WCT normative sample, those with probable dementia of the Alzheimer’s type—mild severity were excluded with the rationale that “previous research has demonstrated that individuals with severe cognitive limitations perform poorly on effort measures” (Pearson, 2009, p. 77). This assertion is disputable, as the use of PVTs has previously been validated in cognitively impaired samples. For example, the effort indices on the Green’s Word Memory Test (WMT; Green, Allen, & Astner, 1996) have been shown to be fairly robust, even in the presence of cognitive impairment (CI; Allen & Green, 1999; Green & Allen, 1999; Green, Iverson, & Allen, 1999). The validation of PVTs in samples with CI of diverse etiologies is important, as suboptimal test engagement and/or exaggeration of symptoms is not solely observed in samples without neurological disorders. Indeed, those with neurological disorders and/or with CI may also have motivation to present with more severe impairment for some external gain, such as increased attention or monetary benefits (Boone, Lu, & Herzberg, 2002b). Thus, clinically useful PVTs should ideally be robust to CI while simultaneously remaining sensitive to patients’ level of task engagement.
In its consensus position article regarding assessment of effort, response bias, and malingering, the American Academy of Clinical Neuropsychology advocated that,
clinicians need to assign weight to specific results according to the rigor of the studies and the relevance of samples studied to the clinician’s case at hand . . . with greater weight being given to indicators that have proven validity across multiple studies. (Heilbronner et al., 2009, p. 1101)
To date, few researchers have independently examined the utility of the WCT in distinguishing valid from invalid performance, particularly in clinical samples, and no universally accepted cut score has been established across diagnostic groups. There have been few published validation studies of the WCT outside of the initial validation sample. Miller et al. (2011) examined WCT performance, both independently and in combination with other measures in the ACS package associated with the Wechsler Memory Scale–Fourth edition (Wechsler, 2009), in a sample of individuals with moderate to severe TBI compared with healthy adults coached to simulate TBI. Results indicated that the WCT added incremental validity to the assessment of performance validity with the ACS system. As a stand-alone, the WCT showed good ability to distinguish feigners from survivors of TBI (Area under the curve [AUC] = .84). Mean WCT scores were provided for the TBI group (M = 46.5; SD = 4.4) and the simulator group (M = 34.7; SD = 11.8), and a cutoff score of ≤41 was found to maximize classification accuracy. Two other studies utilizing a WCT cut score corresponding to a <10% base rate (per the ACS manual) demonstrated similar sensitivity/specificity values for classifying invalid performance (41%/84% in a TBI vs. simulator sample [Bashem et al., 2014]; 38%/96% in a mixed clinical sample [Davis, 2014]). Authors of both studies noted that their findings suggest that the WCT should not be used alone as an indicator of performance validity, but should be considered in the context of performance on multiple PVTs. In a later analogue sample study, Barhon, Batchelor, Meares, Chekaluk, and Shores (2015) compared the WCT against the Test of Memory Malingering (TOMM) in its ability to distinguish between undergraduate students instructed to provide full effort and those instructed to feign acquired brain injury (with and without coaching). Both the TOMM and WCT demonstrated good ability to discriminate between feigned CI and valid performance. A WCT cut score of 42/50 maximized sensitivity (83%) and specificity (96%) in this sample, while use of the recommended TOMM cut score of 45/50 on Trial 1 or Trial 2 was associated with 89% sensitivity and 96% specificity. The measures showed statistically equivalent ability to distinguish feigned CI from valid performance in this sample.
While the ability of the WCT to differentiate noncredible from valid performance in individuals with TBI via simulator studies is encouraging, further validation with other clinical populations is necessary. In particular, it is essential that the utility of the WCT be further examined in mixed clinical samples, with patients exhibiting both valid and invalid performance as assessed with a concurrently administered, well-validated measure of performance validity, such as the WMT (Green, 2003; Green et al., 1996; Green et al., 1999). Specifically, it is important to assess the WCT’s psychometric properties among those both with and without cognitive deficits to determine whether it is able to clearly identify invalid performance even when CI is present. Accordingly, the current study had the following aims: (a) to examine the classification accuracy, sensitivity, and specificity of the WCT for identifying invalid performance in a mixed clinical sample; (b) to compare the classification accuracy, sensitivity, and specificity of the WCT in participants with and without CI; and (c) to establish optimal WCT cut scores for valid/invalid performance. Finally, the TOMM was also included in this study as a comparison measure given its previously established sensitivity to noncredible performance (e.g., Jelicic, Ceunen, Peters, & Merckelbach, 2011; Teichner & Wagner, 2004; Tombaugh, 1996) and status as the most commonly used PVT among clinical neuropsychologists in general (Rabin, Barr, & Burton, 2005; Sharland & Gfeller, 2007) and among neuropsychologists practicing in VA medical centers (Young, Roper, & Arentsen, 2016).
Method
Participants
Data for this retrospective, cross-sectional study were extracted from a mixed clinical sample of clinically referred veterans who completed comprehensive neuropsychological evaluation at a VA Medical Center from 2015 to 2016. The study was approved by the local institutional review board and all participants provided written informed consent to be included in the study after completion of evaluation. Participants who completed the three primary PVTs (i.e., WMT, TOMM, WCT) were eligible for inclusion. Of the 92 initial participants, two were excluded for missing data. Thus, the final sample consisted of 90 veterans (87.8% male; N = 79) with an average age of 54.03 years (SD = 13.46; range = 24-77 years) and average education of 13.78 years (SD = 2.33; range = 7-19 years). The sample was 48.9% Caucasian (N = 44), 35.6% Hispanic (N = 32), 11.1% African American (N = 10), and 4.4% Other (N = 4). Twenty-six percent (N = 24) of the sample was bilingual (English/Spanish). Forty-three percent of the sample (N = 39) obtained a valid WMT score and an additional 21% (N = 19) had valid WMTs with the genuine memory impairment profile (GMIP) and concurrent CI (see Measures section for description of the GMIP calculation). Thus, in total, 64% (N = 58) of the sample had valid WMT performance, whereas 36% (N = 32) were classified as invalid based on the WMT. There were no significant group differences between participants with valid and invalid WMT scores in terms of age, F(1, 88) = 2.052, p = .156, education, F(1, 88) = .011, p = .917, race/ethnicity, χ2(3, 90) = 1.126, p = .771, or language, χ2(1, 90) = .533, p = .465. Neuropsychological test batteries and accompanying norms for each patient were selected by board-certified neuropsychologists on the basis of referral question and patient demographics. Though batteries varied, each included measures to address the domains of processing speed, attention, language, visuospatial skills, memory, and executive functioning. Clinical diagnoses of neurocognitive disorders were determined by these neuropsychologists at the time of neuropsychological evaluation per formal Diagnostic and Statistical Manual of Mental Disorders–Fifth Edition (APA, 2013) diagnostic criteria. Of those whose performance on neurocognitive testing was considered valid on the basis of WMT performance (N = 58), 48.3% (N = 28) were cognitively unimpaired (i.e., did not meet criteria for a major or mild neurocognitive disorder) and 51.7% (N = 30) were cognitively impaired (i.e., met criteria for a neurocognitive disorder). Of those with CI, 80% had a mild neurocognitive disorder (N = 24), while 20% (N = 6) met criteria for major neurocognitive disorder. The most common etiologies were vascular and vascular with comorbid psychiatric diagnosis (N = 13, 43.3%); multiple etiologies, not including vascular (N = 4, 13.3%); other neurological disorders (i.e., epilepsy, frontotemporal lobar degeneration, and Parkinson’s disease; N = 4, 13.3%); Alzheimer’s disease (N = 3, 10%); and severe TBI (N = 2, 6.7%). Additionally, there was one participant (1.1%) with each of the following etiologies: substance-induced, posttraumatic stress disorder, attention-deficit/hyperactivity disorder, and unspecified etiology. Cognitively impaired participants (M = 59.40 years; SD = 13.83) were, on average, approximately 8 years older than the unimpaired participants (M = 51.39 years; SD = 14.63), F(1, 56) = 4.592, p = .036, but otherwise did not differ significantly with respect to education, F(1, 56) = 2.767, p = .102, race/ethnicity, χ2(3, 90) = 1.404, p = .705, or language, χ2(1, 90) = .217, p = .641.
Measures
The following three PVTs were administered in the context of a larger clinical test battery. PVT administration was counterbalanced to minimize order effects.
Word Memory Test
On the basis of previous research demonstrating its robustness for detecting invalid/noncredible test performance with minimal undue impact from external factors (i.e., reading ability, low education and/or intelligence, age, and pain; see Strauss, Sherman, & Spreen, 2006), the WMT was selected as the criterion measure to establish validity groups for this study. The WMT is a computer-administered PVT involving two learning trials in which examinees are presented with 20 word pairs at a rate of one pair every 6 seconds (with a total administration time of approximately 2 minutes), followed by a series of immediate and delayed recall tasks designed to assess task engagement (Green, 2003; Green et al., 1996; Green et al., 1999). The entire test takes a total of 40 to 45 minutes to administer, which includes 10 to 15 minutes of examinees completing subtests with a 30-minute delay in between the immediate and delayed subtests, during which other tests can be administered. It yields three primary effort indices, each with a maximum score of 100% raw percentage correct: Immediate Recognition (IR), Delayed Recognition (DR), and Consistency (CNS). Task failure for each primary effort index is defined as ≤82.5% raw percentage correct (Green et al., 1996), and overall invalid performance is defined by failure on any one of the three primary effort indices. The WMT also includes two “easy” memory indices (Multiple Choice [MC] and Paired Associates [PA]) as well as one “hard” memory index (Free Recall [FR]). In order to avoid misclassification of failure in individuals with genuine memory impairments, an additional index, the GMIP, was developed (see Green, 2003, 2005). Per GMIP guidelines, a difference score of 30 or greater between mean percentage correct for the effort indices and mean percentage correct for the memory indices is indicative of adequate task engagement in the context of genuine memory deficits (Green, Montijo, & Brockhaus, 2011). For the purposes of this study, participant performance was defined as valid if (a) primary effort index scores were all above established cutoff of 82.5% or (b) the GMIP was indicative of adequate task engagement with concurrent CI. Participants with one or more primary effort index scores below established cutoff and a GMIP indicative of inadequate test engagement were classified as having invalid performance.
Advanced Clinical Solutions: Word Choice Test
The WCT (Pearson, 2009) is described as the “single external validity measure” available in the ACS system for detection of suboptimal effort (Pearson, 2009, p. 75). It consists of 50 learning items and 50 test items. During learning trials, 50 words are sequentially shown and read aloud to examinees. In order to increase attention to the words, examinees are asked to classify each word as being either “man-made” or “natural.” Immediately after the learning trials are completed, examinees are presented with a series of 50 word pairs, each containing a target item and a distractor, and are asked to identify the target item. Test administration time is approximately 5 to 10 minutes. The WCT total score is derived from the number of correct responses, with a maximum score of 50.
Test of Memory Malingering
The TOMM (Tombaugh, 1996) is a PVT in which examinees are sequentially presented with a series of 50 simple line drawings for 3 seconds each. Examinees are then presented with 50 pairs of images, each containing one target image and one nontarget image, and are asked to identify target stimuli. This procedure is completed twice (Trial 1 and Trial 2), with an average administration time of 15 minutes for Trials 1 and 2. Following a 15-minute delay, a Retention Trial is administered in which patients must select target items from 50 stimulus pairs presented sequentially. The maximum total score for each trial is 50. In order to provide the fairest and most direct comparison against the single-trial WCT with respect to administration time and number of items, only Trial 1 performance on the TOMM was considered in the current study. Previous research has demonstrated that valid performance on TOMM Trial 1 is highly predictive of valid performance on Trial 2 and the Retention Trial, leading many to argue for a discontinuation after a passed Trial 1 (e.g., Hilsabeck, Gordon, Hietpas-Wilson, & Zartman, 2011; Mossman, Wygant, Gervais, & Hart, 2017; O’Bryant, Engel, Kleiner, Vasterling, & Black, 2007; O’Bryant et al., 2008). In an outpatient VA sample, application of a TOMM Trial 1 cutoff score of ≤40 was associated with 92% classification accuracy (Denning, 2012).
Data Analysis
Correlational analyses were conducted to assess the relationships between the WMT, WCT, and TOMM. Logistic regression was performed to analyze the classification accuracy of the WCT and TOMM in predicting valid versus invalid performance on the WMT separately and in a combined model to determine the independent clinical utility of these measures and the incremental utility of using both measures. Positive/negative predictive values and diagnostic odds ratios (DORs) also were calculated. Finally, receiver operating characteristic (ROC) curve analyses were used to examine the sensitivities and specificities associated with the full range of possible WCT and TOMM scores, and to select cut scores that maximized sensitivity and specificity for identifying both valid as well as invalid performance. By establishing cut scores for identifying invalid and valid performance separately, each measure’s equivocal range of scores (those falling between the valid and invalid cut scores) was elucidated. Following these analyses, the sample was further divided by CI status and analyses were repeated for each group (i.e., cognitively impaired and unimpaired) to evaluate potential differences in WCT and TOMM classification accuracy and/or sensitivity/specificity based on CI status.
Results
Descriptive statistics for each PVT and correlations among the PVTs for the valid participants (N = 58) are presented in Table 1. Performance on the WCT was not significantly correlated with TOMM performance, suggesting that the two measures tap differing constructs. An interesting pattern of divergence between the WCT and TOMM also emerged when examining their correlations to WMT indices. Most notably, while the WCT was significantly correlated with the WMT primary effort indices (i.e., IR, DR, CNS), it also demonstrated robust correlations with the WMT subtests that are more reliant on actual memory function (i.e., MC, PA, FR), whereas the TOMM was significantly associated with the WMT primary effort indices, but had small correlations with the “easy” memory subtests (i.e., MC and PA) and nonsignificant correlation with the “hard” memory subtest (i.e., FR).
Performance Validity Test (PVT) Performance and Correlations.
Note. WMT = Word Memory Test; IR = Immediate Recall; DR = Delayed Recall; CNS = Consistency; MC = Multiple Choice; PA = Paired Associates; FR = Free Recall; TOMM = Trial 1 of the Test of Memory Malingering; WCT = Word Choice Test.
Correlations were computed only for patients with valid performance on the WMT (N = 58).
p < .05. **p < .01. ***p < .001.
Logistic regression was performed to examine classification accuracy of the WCT and TOMM for predicting validity group membership based on WMT scores. As noted in Table 2, when the entire sample was examined, both the WCT and TOMM independently emerged as significant predictors with overall classification accuracies of 80% and 86.7%, respectively. When both the WCT and TOMM were entered into the same model, each PVT remained a significant predictor variable, and overall correct classification improved by 2.2% to 88.9%. Next, the sample was divided by CI status to allow for further examination of the classification accuracy of the WCT and TOMM for discriminating valid from invalid performance. Among those without CI (see Table 3), the WCT and TOMM performed similarly with respect to classification of performance validity, with each exhibiting at least 85% accuracy. When both measures were entered into the logistic regression model, each remained a significant predictor, and the overall classification accuracy rate improved to 90%. In contrast, a different pattern emerged for those with CI (see Table 4). Namely, both the WCT and TOMM independently were significant predictors; however, the WCT’s overall classification accuracy to discriminate valid from invalid performance decreased to 69% (relative to 85% accuracy among cognitively unimpaired participants), whereas the TOMM showed much less decrease and maintained 82% accuracy (compared with 86.7% accuracy among cognitively unimpaired participants). The TOMM’s DOR also remained well above 20, which is consistent with potentially useful tests (Fisher, Bachmann, & Jaeschke, 2003), whereas the WCT’s DOR notably decreased. Additionally, when both measures were entered into the logistic regression model, overall classification accuracy was approximately 84%, but the WCT became nonsignificant, whereas the TOMM remained a significant predictor.
Logistic Regression Analyses Predicting WMT Validity Group Classification for Invalid (N = 32) and All Valid Participants (N = 58).
Note. TOMM = Test of Memory Malingering; WCT = Word Choice Test; PPV = Positive Predictive Value for identifying invalid performance; NPV = Negative Predictive Value for identifying invalid performance; DOR = diagnostic odds ratio; WMT = Word Memory Test; SE = standard error.
Sensitivity refers to identification of invalid performance.
Logistic Regression Analyses Predicting WMT Validity Group Classification for Invalid (N = 32) and Cognitively Unimpaired Valid Participants (N = 28).
Note. TOMM = Test of Memory Malingering; WCT = Word Choice Test; PPV = Positive Predictive Value for identifying invalid performance; NPV = Negative Predictive Value for identifying invalid performance; DOR = diagnostic odds ratio; WMT = Word Memory Test; SE = standard error.
Sensitivity refers to identification of invalid performance.
Logistic Regression Analyses Predicting WMT Validity Group Classification for Invalid (N = 32) and Cognitively Impaired Valid Participants (N = 30).
Note. TOMM = Test of Memory Malingering; WCT = Word Choice Test; PPV = Positive Predictive Value for identifying invalid performance; NPV = Negative Predictive Value for identifying invalid performance; DOR = diagnostic odds ratio; WMT = Word Memory Test; SE = standard error.
Sensitivity refers to identification of invalid performance.
ROC curve analyses were utilized to determine appropriate cutoffs to maximize sensitivity and specificity for the WCT and TOMM. For each measure, two ROCs were constructed, one for predicting valid performance and one for invalid performance as determined by the WMT. AUCs and sensitivity/specificity values for the overall sample, as well as those with and without CI, are presented in Table 5. In sum, TOMM AUCs and cut scores were essentially stable across groups, with similar sensitivity and specificity for predicting both valid and invalid performance for both those with and without CI. Moreover, the TOMM demonstrated a robust ability to maximize both sensitivity (i.e., ≥78%) and specificity (i.e., ≥87%) across groups. In contrast, the WCT showed a more striking sensitivity/specificity trade-off that was most pronounced for those with CI. For example, among the cognitively impaired group, a TOMM cut score of 44/45 was associated with a sensitivity of 83% and specificity of approximately 94% for identifying valid performance. In order for the WCT to maintain a similar level of specificity (i.e., >90%), a cut score of 47/48 would be selected and resulted in a 40% decrease in sensitivity to 43%. Likewise, despite equal specificities (i.e., 87%), the WCT’s sensitivity (i.e., 56%) for identifying invalid performance also was considerably less than the TOMM’s (i.e., 78%). Finally, examination of each measure’s optimal cut scores revealed equivocal zones (i.e., the range of scores that fell between the optimal cut score for identifying valid performance and the optimal score for identifying noncredible performance, and are thus not clearly indicative of either valid or invalid performance) for both measures. Specifically, this was a TOMM score of 44 and WCT scores of 45 to 46 for unimpaired participants, and TOMM scores of 41 to 44 and WCT scores of 43 to 46 among impaired participants.
Receiver Operating Characteristic (ROC) Curve Analyses for All Valid Participants, Cognitively Unimpaired Valid Participants, and Cognitively Impaired Valid Participants.
Note. TOMM = Test of Memory Malingering (Trial 1); WCT = Word Choice Test; AUC = Area under the curve. All ROC curve analyses compared valid performance against invalid participants (N = 32).
Sensitivity refers to identification of valid performance. bSensitivity refers to identification of invalid performance.
p < .001.
Discussion
The WCT is a relatively recent contribution to the growing catalog of PVTs. In its initial validation, performance base rates were examined in a mixed clinical sample (Pearson, 2009); however, no cut score was provided at which performance should be considered invalid. Furthermore, WCT performance was not validated against concurrent, well-validated measures of performance validity. To date, there has been limited examination of the WCT’s validity in diverse clinical samples. The current study was intended to provide further clinical validation for the WCT and its ability to differentiate valid from invalid performance in a mixed clinical sample of patients both with and without CI. An ideal PVT should be sensitive to effort/performance validity and robust to the effects of CI. In the current sample, the TOMM demonstrated such durability, as its sensitivity and specificity varied little between groups with and without CI. Indeed, its specificity was nearly identical across clinical groups, indicating that the presence of CI did not increase the likelihood of misclassifying valid performance as suboptimal. By contrast, the sensitivity of the WCT differed substantially for groups differing in CI status. To reach acceptable levels of specificity among patients with CI (ideally 90% or greater; Larrabee & Berry, 2007), there was a steep sensitivity trade-off, such that clinicians would have essentially a 50/50 chance of classifying invalid performance as such. Alternatively, acceptable sensitivity could be achieved only at the cost of greatly increased risk of false positive errors (i.e., incorrectly classifying performance as invalid). Thus, while the WCT demonstrated similar properties to the TOMM among the cognitively unimpaired, these results suggest that its clinical utility as a PVT is more limited among those with CI and should be interpreted with greater caution in this population.
Closer examination of the correlations among PVTs sheds some light on why the WCT performed worse in the presence of CI in the current sample. Specifically, the WCT was highly correlated with all six indices of the WMT, including the WMT-FR subtest. Conversely, TOMM Trial 1 performance was strongly correlated with WMT primary effort indices, but only weakly correlated with the “easy” memory indices (i.e., MC,PA) and nonsignificantly correlated with the WMT-FR subtest. Emerging evidence suggests that the WMT-FR index acts as a bona fide measure of memory function (e.g., Armistead-Jehle, Green, Gervais, & Hungerford, 2015; Eichstaedt et al., 2014; Soble et al., 2016). Thus, the pattern of correlations between the WCT and the WMT suggests that the WCT is tapping cognitive processes related to genuine memory abilities, in addition to cognitive processes related to effort/performance validity. By contrast, the lack of correlation between TOMM scores and the WMT-FR subtest, in the context of strong correlations between the TOMM and WMT effort indices, suggests that the TOMM is differentially picking up on cognitive processes related to effort/performance validity with less reliance on cognitive processes related to memory. This is consistent with the very small change in sensitivity/specificity trade-off for the TOMM when comparing cognitively impaired and cognitively unimpaired groups in this sample, relative to the larger change in WCT sensitivity/specificity between the two groups. Comparison of the ROC curves identifying “valid” performance and those identifying “invalid” performance revealed both the TOMM and WCT had an equivocal zone, or range of scores which fall between the two cutoffs in which interpretation of performance as valid or invalid is unclear, that was larger among cognitively impaired participants relative to their unimpaired peers. Notably, in the current sample, the optimal cutoff for identifying invalid performance on the TOMM Trial 1 (<43) was approximately similar to cutoffs reported in previous examinations of its classification accuracy (Denning, 2012).
The current study had several limitations. Because all participants were referred for evaluation, the cognitively unimpaired group was a clinically presenting group with subjective cognitive complaints, rather than a traditional control group. Outside of the target PVTs, test batteries used to determine the presence or absence of CI were not identical across participants, but rather were individualized to the patient. Another limitation was the use of retrospective clinical data, which limits the generalizability of results and introduces the complication of sampling bias. Notably, data were restricted to those who were able to tolerate a full evaluation and gave informed consent for data collection, which may limit generalizability. Third, the use of WMT as a single criterion measure of performance validity is a limitation. While the WMT has excellent psychometric properties and has been validated in multiple clinical samples, like all PVTs, it has its own false positive rate, though the GMIP aims to minimize such misclassification (Green et al., 2011). Further research examining the WCT’s classification accuracy should be conducted with validity groups defined by failures on multiple well-validated PVTs (Larrabee, 2005). Moreover, our overall rate of PVT failure was approximately 36%, which is similar to the 40% reported base rate of malingering in litigation samples (Larrabee, 2003) as well as prevalence rates that have been reported in clinical evaluations of active duty service members (54%; Armistead-Jehle & Buican, 2012) and in VA clinical (21%; Sawyer, Testa, & Dux, 2017) and survey (23%; Young et al., 2016) research. Thus, examination of the WCT in other clinical samples with varying base rates of PVT failure is warranted. Finally, research has shown differential PVT performance between mild and major neurocognitive impairment (e.g., Teichner & Wagner, 2004) with many PVTs demonstrating much lower specificity among patients with dementia (Dean, Victor, Boone, Philpott, & Hess, 2009). It was not possible to compare those with mild versus major neurocognitive disorder in the current sample due to inadequate sample size, and this is an area for further research. While current findings offer some support for the use of the WCT in a mixed clinical sample, additional validation studies are needed to further establish the measure’s utility and accuracy. Given that the WCT performed differently for patients with and without CI, it would be beneficial for future research to validate its use with clearly defined clinical samples with diverse neurological and psychiatric conditions and ideally to establish cut scores which maximize WCT sensitivity and specificity among these individuals. Such procedures have been utilized in the development of other commonly used PVTs (e.g., b Test, Boone, Lu, & Herzberg, 2002a; Dot Counting Test, Boone et al., 2002b), increasing examiners’ ability to gauge patients’ performance validity against normative groups that share important clinical characteristics.
Despite its limitations, this study contributes to the flourishing literature which highlights the vital importance of considering a PVT’s validation samples when interpreting results. Continued research in this area, examining the robustness of PVTs to CI and the utility of PVTs with diverse clinical samples, is essential as neuropsychologists answer the call to assess performance validity with multiple measures which are well-validated in the populations which they serve.
Footnotes
Acknowledgements
We would like to acknowledge Justin O’Rourke, PhD, ABPP, and Octavio Santos, MS, for reviewing an initial draft of this article.
Authors’ Note
The views expressed herein are those of the authors and do not necessarily reflect the views or the official policy of the U. S. Department of Veterans Affairs or the U.S. Government.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
