Abstract
Introduction
It is now acknowledged that ADHD is not just a childhood disorder. Rates of persistency ranging from 20% to more than 80% show that it is also present in adulthood with a prevalence rate between 4% and 5% (Barbaresi et al., 2013; Davidson, 2008; Kooij et al., 2010). Untreated, it often causes unfavorable impairments at school; social, occupational, and interpersonal difficulties; and a high number of comorbid disorders like depression, conduct disorders, substance abuse, and delinquency.
An ADHD diagnosis implies certain advantages. Academic benefits include, for example, additional time to complete assignments, no spelling penalties, privileged seating in classrooms, calm testing environments, and others (Sansone & Sansone, 2011). The prescription of stimulant medication, which can also be regarded as an incentive, might be more associated with misuse, diversion, and malingering than previously thought (Rabiner, 2013). In a U.S. representative sample comprised of 18- to 49-year-olds, almost 20% admitted that they received this medication by feigning the respective symptoms (Clemow & Walker, 2014; Novak, Kroutil, Williams, & Van Brunt, 2007). The most commonly reported motives were to enhance cognitive performance and recreational reasons, although data on these supposed positive effects do not exist. Non-medical users of stimulants are more likely to consume alcohol, cigarettes, and other drugs (Rabiner, 2013). Therefore, adolescents and adults undergoing diagnostic evaluations might be motivated to exaggerate symptoms on self-report measures and neuropsychological tests. This highlights the importance of identifying those who feign the symptoms of ADHD to obtain prescriptions for stimulant medication.
It is estimated that up to 40% of individuals who might economically benefit by receiving a particular diagnosis may be responding to neuropsychological tests in a non-credible way (Alfano & Boone, 2007). Probable malingering in psychoeducational contexts was estimated to be around 10% (Pella, Hill, Shelton, Elliott, & Gouvier, 2012). In contrast, the possibility of malingering in the context of ADHD has long been ignored (Alfano & Boone, 2007).
Malingering students could not be differentiated on an ADHD self-rating scale from a control group or a group of students diagnosed with ADHD (Quinn, 2003). They were distinguished from the other groups in a Continuous Performance Test (CPT) as they scored significantly worse on several subtests. In another study (Suhr, Sullivan, & Rodriguez, 2011), differentiation via a CPT was more difficult to identify as the mean scores of the non-credible group were not much different from the ADHD group. After applying clinical criteria (number of CPT subtests with T-scores ≥60), the malingering group performed significantly worse.
Instructed simulators could be detected by high scores on the Conners Adult ADHD Rating Scales (CAARS) and low scores in reading fluency and processing speed subtests but with an error rate of 25% (Harrison, Edwards, & Parker, 2007). The authors argue against the exclusive use of self-rating inventories for diagnosing adult ADHD. Young and Gross (2011) proposed to apply the Minnesota Multiphasic Personality Inventory Test (MMPI-2) in adult ADHD assessment as they were able to differentiate a malingering group and an ADHD group by a number of integrated validity scales in their study (Young & Gross, 2011). Using the ADHD Current and Childhood Symptoms Scales, there were no significant differences between individuals trying to feign symptoms of ADHD and an ADHD clinical group. These results reveal that neuropsychological and self-rating measures do not reliably differentiate between malingerers and those with a true diagnosis of ADHD.
Frazier, Frazier, Busch, Kerwood, and Demaree (2008) conducted a study to detect simulated ADHD and reading disorder using symptom validity measures. Some scores of the Validity Indicator Profile (VIP) and the Victoria Symptom Validity Test (VSVT) were able to detect sub-optimal effort in ADHD. A limiting factor was that the ADHD sample consisted of undergraduate introductory psychology students who were instructed simulators. Other studies also hint at the possibility that symptom validity tests might be better able to differentiate between the two groups.
Lee Booksh, Pella, Singh, and Drew Gouvier (2010) compared undergraduate students simulating ADHD with both a control group and students diagnosed with ADHD by objective measures of attention, symptom validity tests, and self-report measures. The ADHD simulators performed worst on the objective neuropsychological tests, especially on the response time variability index of the CPT. Simulators were also able to feign symptoms on self-report measures, making them appear similar to the clinical group. Based on data from the Word Memory Test (WMT), a well-examined symptom validity test, 42% of the simulators were misclassified as controls. This is a disappointing finding, but further studies applying symptom validity tests produced more stringent results. Suhr, Hammers, Dobbins-Buckland, Zimak, and Hughes (2008) cited studies showing a relatively high base rate of ADHD-like symptoms in the general population and in individuals seeking treatment due to other diagnoses (Suhr et al., 2008). The authors conclude that it is problematic to rely on self-report information for the diagnosis of ADHD in adults. In their own study, their sample consisted of individuals who had undergone neuropsychological evaluations in a university psychology clinic. They compared those who failed the WMT, which were 31% of the total sample, with a group of ADHD patients and a group with other psychological symptoms. The three groups scored not significantly different on the different CAARS scales. The invalid group and the ADHD group were not significantly different on the Wender Utah Rating Scale (WURS-k) of childhood ADHD symptoms. The non-credible group was significantly worse than the other groups in memory performance and executive functioning. Contrary to this, an association between symptom exaggeration and non-credible performance was found in a different study. In a sample of college students referred for ADHD evaluation, 47.6% failed one or more effort measures of the WMT (Sullivan, May, & Galbally, 2007). Those who failed the WMT had lower performance IQ scores and higher scores on the CAARS. The authors conclude that WMT failure is an indicator of symptom exaggeration and negative response bias.
Marshall et al. (2010) examined 268 adults referred for ADHD assessment with symptom validity measures to identify symptom exaggeration. In their sample, 22% were suspected of showing negative response bias. Several indices of different symptom validity tests showed acceptable sensitivity and specificity values. Feigning behavior occurred most frequently in symptom validity tests, which seemed to measure attention. It was impossible to detect non-credible performance with neuropsychological tests or self-report scales.
Jasinski et al. (2011) instructed undergraduates with a history of ADHD to respond honestly or to exaggerate symptoms and compared them with undergraduates with no history of ADHD who were also instructed to respond honestly or to feign symptoms of ADHD (Jasinski et al., 2011). They argue for the application of at least two symptom validity tests that resulted in a sensitivity of .475 and a specificity of 1.00 in their study.
In their review, Musso and Gouvier (2014) concluded that no available ADHD self-report questionnaire was robust against simulation so that a larger amount of false positive diagnoses could result if one relied solely on such information. Embedded validity indices were unable to detect malingered ADHD. In neuropsychological tests, the profiles of ADHD simulators were relatively similar to those with a diagnosis of ADHD. The most promising approach to detect malingered ADHD might be the use of several symptom validity tests. They possess excellent specificity but only moderate sensitivity. Tucha, Fuermaier, Koerts, Groen, and Thome (2014) reached similar conclusions, but they were more pessimistic in that they stated that new measures and approaches to detect feigned ADHD are missing.
In the current study, we analyzed data from patients who were diagnosed with ADHD. All patients received an ADHD diagnosis by experienced and licensed clinical psychologists using a comprehensive diagnostic strategy. There were also patients who failed a symptom validity test and still received the diagnosis because other diagnostic information justified this and because failure on a symptom validity test does not automatically indicate that an individual is malingering (Suhr et al., 2008). We compared those who failed a symptom validity test with those who passed the test to explore whether there are signs of negative response bias on group level.
Method
Participants
In our outpatient department, 292 patients were consecutively seeking diagnostic counseling regarding ADHD from February 2011 to February 2015. Fifty-eight (19.8%) participants received no ADHD diagnosis, and 38 (13%) discontinued the diagnostic process. Therefore, a study sample of 196 (67%) adult patients resulted (Figure 1).

Emergence of the study sample.
In our study sample, there were 127 men (65%) with a mean age of 32.1 years (SD = 10.7, range = 18-62 years) and 69 women (35%) with a mean age of 33.2 years (SD = 11.0, range = 18-59 years). They were diagnosed with ADHD by experienced licensed clinical psychologists on the basis of a detailed clinical history, the Wender-Reimherr Interview (WRI), the Conners Adult ADHD Rating Scales–Long version/Self-Rating (CAARS-L: S) and Observer-Rating (CAARS-L: O), the WURS-k, the German ADHD Self-Rating Scale (ADHS-SB), the Quantified Behavior Test Plus (Qb+©), and the Test of Attentional Performance (TAP: GO/NOGO, Divided Attention, Sustained Attention). In addition, the Amsterdam Short Term Memory Test (AKGT) was also applied as a symptom validity measure. All participants gave written informed consent. Our study conforms to the Declaration of Helsinki and was approved by the local ethics committee of the Faculty of Psychology at the Philipps University in Marburg, Germany.
Materials
Wender-Reimherr Adult Attention Deficits Disorders Scale (WRAADDS)
The German version of the WRAADDS is a psychopathological interview that measures 28 psychological characteristics in seven symptom areas: attention deficits, hyperactivity, temper, affective lability, emotional overreaction, disorganization, and impulsivity (Corbisiero, Buchli-Kammermann, & Stieglitz, 2010). We did not use it for quantitative analyses, but as one component, together with the clinical history and the other measures, to establish or refute the diagnosis. Cutoff scores for the different symptom domains are as follows: inattention ≥ 5, hyperactivity ≥ 3, temperament ≥ 3, affective lability ≥ 4, emotional overreaction ≥ 5, impulsivity ≥ 5, global impairment per symptom domain ≥ 2 (Rösler et al., 2007).
CAARS-L: S and CAARS-L: O
The German version of the CAARS-L: S assesses ADHD symptoms in adults aged 18 years or older. Symptoms are rated on a Likert-type scale (0 = not at all/never to 3 = very much/very frequently). The long version consists of 66 items, but only 42 items were included in the original factor analysis by Conners, Erhardt, and Sparrow (1999) due to statistical restrictions made by the authors. Four factors emerged from their analyses: inattention/memory problems, hyperactivity/restlessness, impulsivity/emotional lability, and problems with self-concept. Confirmatory factor analyses of the German version in healthy adults and ADHD patients supported this factor analytic solution (Christiansen et al., 2013; Christiansen et al., 2011). The four subscales were significantly influenced by age, gender, and the number of years in education. Symptom severity decreased with increasing age, males scored higher than females on hyperactivity and sensation-seeking behavior, and females scored higher than males on problems with self-concept. Overall symptom ratings were higher for individuals who had received less education. Test–retest reliability ranged between .85 and .92; sensitivity and specificity were high for all four subscales. The CAARS-L: S represents a reliable and cross-culturally valid measure of current ADHD symptoms in adults (Christiansen et al., 2012). The same is true for the observer version of the scale (CAARS-L: O), in which a person who has a close relationship to the participant under examination should rate the same items (Christiansen, Hirsch, Abdel-Hamid, & Kis, 2014). The hypothesized factor structure could also be supported, and the observer version also possesses good statistical quality criteria. The maximum scores for the subscales are as follows: Inattention/Memory Problems = 36, Hyperactivity/Restlessness = 36, Impulsivity/Emotional Lability = 36, Problems With Self-Concept = 18, Diagnostic and Statistical Manual of Mental Disorders (4th ed.; DSM-IV; American Psychiatric Association, 1994) Inattention = 27, DSM-IV Hyperactivity/Restlessness = 27, DSM-IV Total Symptom Score = 54, DSM-IV ADHD Index = 36. Norms based on gender and age are provided as well as T-scores (M = 50, SD = 10), with participant scoring ≥65 representing the 85th percentile and above (Christiansen et al., 2014).
The Inconsistency Index was created as a validity measure to assess response inconsistency (Conners et al., 1999). The long version of the CAARS contains eight pairs of items with similar content. The absolute value of the differences between each item pair is calculated, and these differences are added together. A cutoff score of eight or greater is said to represent non-credible responding.
The other validity index used in this study was introduced by Suhr, Buelow, and Riddle (2011) and is called the Infrequency Index. For this, 12 items from the scales Inattention/Memory Problems, Impulsivity/Emotional Lability, and DSM-IV Hyperactive/Impulsive Symptoms are selected that occur “pretty much, often” to “very much, very frequently” by 10% or less of the total participant sample. In their empirical analyses, a score of 21 or higher was an indicator for non-credible performance; this was also used in our study.
WURS-k
The German version of the WURS-k (Retz-Junginger et al., 2003; Retz-Junginger et al., 2002) retrospectively assesses ADHD-relevant childhood behaviors and symptoms in adults. It consists of 21 items that distinguish patients with ADHD from a non-patient comparison group. Participants are instructed to rate 25 items (4 are control items) that complete sentence stems such as “As a child I was or had . . . .” Ratings are to be completed on a 5-point Likert-type scale (0 = not at all or very slightly to 4 = very much). The maximum score is 84. Test–retest reliability and Cronbach’s alpha were around .90. Factor analyses generated a five-factor solution with the factors inattention/hyperactivity, impulsivity, anxiety/depression, oppositional behavior, and social adaptation by using 21 items. A total score ≥30 indicates the possibility of ADHD during childhood.
ADHS-SB
The ADHS-SB consists of the 18 DSM-IV items that are broken down into the factors “inattention” (9 items), “hyperactivity,” and “impulsivity” (9 items together; Rösler et al., 2004). The items are scored on a 4-point Likert-type scale (0 = not at all to 3 = very pronounced/almost always the case). The maximum score is 54. Test–retest reliability coefficients were between .78 and .89. Correlations with subscales of the NEO Five Factor Inventory were in the expected directions. A total score ≥18 indicates the possibility of adult ADHD.
Qb+©
The computerized Qb+© is a combined continuous performance and activity test for participants 12 years and older using a simultaneous high resolution motion tracking system. It separately assesses hyperactivity, inattention, and impulsivity and lasts 20 min. Presented stimuli are a blue circle, a blue square, a red circle, and a red square. A response key should be pressed when two identical stimuli are shown in succession. The ratio of target to non-target stimuli is 25:75. Normative data have been gathered from 1,307 individuals between 6 and 60 years with an even age and gender distribution (Ulberstad, 2012). Q-scores based on age and gender are derived for hyperactivity, inattention, and impulsivity. A Q-score ≥1.5 is regarded as an atypical result. Psychometric properties with respect to sensitivity (86%) and specificity (83%) have recently been published (Edebol, Helldin, & Norlander, 2013).
TAP (GO/NOGO, Divided Attention, Sustained Attention)
The computer-based test battery for the assessment of attention (TAP) was developed by Zimmermann and Fimm (2012) and measures several dimensions of attention. For our study, we chose the subtests GO/NOGO, Divided Attention, and Sustained Attention.
GO/NOGO measures selective attention. In the test form “2 of 5” (2 critical stimuli among 5 stimuli), a sequence of five squares with different patterns appears on the screen. Two of these squares are defined as target stimuli, and the patient has to press a special response key as fast as possible when they appear. In Divided Attention, a visual and an auditory task must be processed in parallel. The visual task consists of a quadratic region in which a varying number of crosses appear simultaneously. When four of these crosses form a square, the response key should be pressed as quickly as possible. The auditory task is composed of high and low tones in sequence. When the same tone occurs twice, then the response key should be pressed as quickly as possible. In Sustained Attention, a sequence of stimuli is presented on the monitor. The stimuli vary in a range of feature dimensions: color, shape, size, and filling. A target stimulus occurs whenever two successive patterns have the same shape.
Several studies show the clinical validity of the TAP. Plohmann et al. (1998) could demonstrate the sensitivity of the TAP to measure changes after the neuropsychological treatment of 22 patients with multiple sclerosis. The authors were able to show that differential interventions are needed for different attentional components.
With hierarchical cluster analysis and multidimensional scaling, Sturm, Hartje, Orgass, and Willmes (1993) revealed the construct validity of the TAP. Sprengelmeyer, Lange, and Hömberg (1995) could separate patients with Huntington’s disease from controls based on their TAP results.
We selected the parameters reaction time variability and omission errors as these were shown to represent stable features of ADHD (Adams, Roberts, Milich, & Fillmore, 2011; Kofler et al., 2013). T-scores based on age and gender are available.
AKGT
The AKGT measures negative response bias and insufficient motivation in psychological examinations (Schmand & Lindeboom, 2005). It is presented as a test of short term memory and attention. Five semantically related words are shown for 8 s. They should be read aloud and remembered. Then a simple arithmetic problem is given. After this, five words are again presented, three of which were previously shown. Those three should be identified by the participant. Thirty tasks total a maximum of 90 points. The reliability of the test is satisfactory. Internal consistency in different samples was around .90. In a sample of mixed neurological patients, test–retest correlation was .85 within an interval of 1 to 3 days. The test also shows good validity. The cutoff value for the AKGT is ≤84 points. Sensitivity for lack of motivation was 91% (in experimental simulants) and specificity 89% (in neurological patients). Healthy controls from age 9 on master this test almost perfectly. Patients with neurological disorders like concussion, brain tumors, multiple sclerosis, or difficult-to-treat epilepsy rarely have difficulties in handling this test, provided that they do not have serious cognitive deficits.
The purpose of our study was to examine the rate of failure in a symptom validity measure in adults diagnosed with ADHD. We compared those patients who scored below the cutoff level of the AKGT (≤84 points) with those who passed the test.
Statistical Analyses
Quantitative variables were first analyzed with descriptive statistics. When standard deviations were very close to or greater than means, Huber’s M estimators were calculated. Comparisons on categorical variables were done with χ2 tests (Sprent & Smeeton, 2007). Differences in continuous variables were tested with t tests, Mann–Whitney U tests, and the multivariate approach of the General Linear Model (GLM; Tabachnick & Fidell, 2007). The size of the resulting effects was examined with the effect sizes probability of superiority (PS) and eta2 (Grissom & Kim, 2012). The effect size “PS estimate” is calculated as follows: PS = U/(n1×n2). The PS indicates the probability that a randomly selected participant of group n1 has a higher score than a randomly selected participant of group n2. A PS of .5 means that both groups are equal regarding a specific variable and that there is no effect. Consequently, the larger the effect, the more the PS deviates from .5. An eta2 of .03 denotes a small effect; a value of .06, a medium effect; and a value of .09 and higher, a large effect. Pearson and Spearman correlations were used to calculate bivariate associations. Multiple testing was corrected with the Bonferroni method. All calculations were performed using IBM SPSS Statistics 22.
Results
Sixty-three patients (32.1%) scored below the cutoff level of the AKGT (≤84 points). Their mean score was 75.4 points (SD = 7.2), compared with 87.9 points (SD = 1.7) in those who passed the test. This difference is statistically significant (Mann–Whitney U test, Z = −11.46, p < .001). The effect size PS has a value of .24 and deviates largely from the chance level at .50, therefore signaling a large effect. The demographic characteristics of both subsamples are displayed in Table 1.
Demographic Characteristics in Those Who Failed the AKGT (n = 63) Compared With Those Who Passed the AKGT (n = 133).
Note. AKGT = Amsterdam Short Term Memory Test.
Both groups did not differ regarding age (t test, p = .08), gender (χ2 test, p = .71), or education (χ2 test, p = .23).
Self-Report of ADHD Symptoms
The two groups do not differ significantly regarding the self-report of ADHD symptoms, Wilks’s Lambda = .93, F(11, 184) = 1.22, p = .28, eta2 = .068.
The cutoff values of the WURS-k (>29 points) and of the ADHS-SB (>17 points) are largely exceeded in both groups (Table 2). The majority of the CAARS-S scales have T-scores ≥65 and <80, regardless of different age and gender norms, indicating clinically relevant symptoms in both groups. The means of the Inconsistency Index are under the recommended cutoff value of ≥8 (Conners et al., 1999); 15.9% of those who failed the AKGT have an Inconsistency Index ≥8 versus 23.3% in those who passed the test, χ2(df = 1) = 1.43, p = .23. The means of the Infrequency Index are under the recommended cutoff value of ≥21 (Suhr, Buelow, et al., 2011); 41.3% in the AKGT fail group have an Infrequency Index ≥21 versus 33.1% in those who passed the test, χ2(df = 1) = 1.25, p = .26.
Means and Standard Deviations (SD) of ADHD Self-Report Measures in Those Who Failed the AKGT (n = 63) Compared With Those Who Passed the AKGT (n = 133).
Note. AKGT = Amsterdam Short Term Memory Test; CAARS = Conners Adult ADHD Rating Scales; DSM-IV = Diagnostic and Statistical Manual of Mental Disorders (4th ed.; American Psychiatric Association, 1994).
CAARS Observer Ratings (CAARS-O)
The two groups do not differ significantly regarding observer ratings of ADHD symptoms, Wilks’s Lambda = .94, F(9, 186) = 1.43, p = .18, eta2 = .065.
The majority of the CAARS observer scales also have T-scores ≥65 and <80, regardless of different age and gender norms, indicating clinically relevant symptoms in both groups (Table 3). The means of the Inconsistency Index are under the recommended cutoff value of ≥8 (Conners et al., 1999); 31.7% of the observers of those who failed the AKGT have an Inconsistency Index ≥8 versus 21.1% of the observers of those who passed the test, χ2(df = 1) = 2.64, p = .10. The means of the Infrequency Index are under the recommended cutoff value of ≥21 (Suhr, Buelow, et al., 2011); 36.5% of the observers of those who failed the AKGT have an Infrequency Index ≥21 versus 27.1% of the observers of those who passed the test, χ2(df = 1) = 1.81, p = .18.
Means and Standard Deviations (SD) of CAARS Observer Ratings in Those Who Failed the AKGT (n = 63) Compared With Those Who Passed the AKGT (n = 133).
Note. CAARS = Conners Adult ADHD Rating Scales; AKGT = Amsterdam Short Term Memory Test; DSM-IV = Diagnostic and Statistical Manual of Mental Disorders (4th ed.; American Psychiatric Association, 1994).
The scales of the CAARS observer version correlate from r = .21 (DSM-IV ADHD Index, p = .02, Pearson) to r = .69 (Inattention, p < .001) with the respective scales in the self-report form (CAARS-L: S) in the two groups. The significance level after Bonferroni correction was .05 / 16 = .003. Higher correlations emerged in the AKGT fail group in the subscale “Inattention” (r = .69 vs. r = .34, p = .002) and in the subscale “DSM-IV Inattentive Symptoms” (r = .60 vs. r = .30, p = .01).
Neuropsychology
The two groups differed significantly regarding neuropsychological parameters, Wilks’s Lambda = .87, F(11, 184) = 2.48, p = .006, eta2 = .129. Those who failed the AKGT had higher reaction time variabilities in selective attention (medium effect size), higher omission errors in selective attention (small effect size), higher reaction time variabilities in Auditory (medium effect size) and Visual divided attention (small effect size), higher omission errors in sustained attention (small effect size), and a more atypical result in the Qb+© subdomain Inattention (small effect size; for details, please see Table 4).
Means and Standard Deviations (SD) of Neuropsychological Parameters in the Test of Attentional Performance (TAP) and the Quantified Behavior Test Plus (Qb+©) in Those Who Failed the AKGT (n = 63) Compared With Those Who Passed the AKGT (n = 133).
Note. AKGT = Amsterdam Short Term Memory Test; M = Huber’s M estimator.
—Could not be calculated because of extremely centralized distribution around the median.
For the GO/NOGO variability, the 95% confidence interval of those who failed the AKGT ranges from T-score 28 to T = 38; for those who passed the AKGT, the T-scores range between 40 and 47. The omission errors in both groups can be considered as average performance (Table 4).
The 95% confidence interval in the Auditory divided attention variability in those who failed the AKGT signals a below-average performance (T = 25 to T = 34) compared with those who passed the test (T = 36 to T = 42). The omission errors in the AKGT fail group are also in the below-average area (T = 29 to T = 33), while the 95% confidence interval in the AKGT pass group includes average performance (T = 33 to T = 45).
Both groups managed to achieve average results in the Visual divided attention variability (AKGT fail: T = 49 to T = 64, AKGT pass: T = 62 to T = 70, 95% confidence intervals) and in Visual divided attention omission errors (AKGT fail: T = 38 to T = 49, AKGT pass: T = 44 to T = 54; 95% confidence intervals).
In the Sustained attention variability, both 95% confidence intervals reach the average area (AKGT fail: T = 36 to T = 42, AKGT pass: T = 36 to T = 43). Regarding Sustained attention omission errors, those who failed the AKGT perform slightly below average (T = 31 to T = 40) compared with those who passed the test (T = 38 to T = 42).
Both groups have atypical mean hyperactivity Q-scores in the Qb+©, while only those who failed the AKGT have a slightly elevated mean inattention score.
Intercorrelations Between Validity Measures
We calculated intercorrelations between the AKGT total score and the CAARS validity indices to explore the nature of these associations and whether there are differences between the two groups.
There are no meaningful correlations between the AKGT total score and any of the CAARS validity indices. Significant associations occur between the Infrequency self-report and Observer Index in both groups. A significant negative correlation between the Infrequency Observer Index and the Inconsistency Observer Index was only found in the group that failed the AKGT. The intercorrelations of the CAARS validity indices are relatively low and do not reach significance. None of the correlations in Table 5 are significantly different between the two groups.
Intercorrelations (Spearman-rho) Between the Total Score of the AKGT and CAARS Validity Indices of the Self-Report and Observer Versions.
Note. Correlations for those who failed the AKGT (n = 63), and for those who passed the AKGT (n = 133) in parenthesis. Correlations printed in bold are significant after Bonferroni correction (p = .05 / 10 = .005). AKGT = Amsterdam Short Term Memory Test; CAARS = Conners Adult ADHD Rating Scales.
Correlations Between Validity Measures and Self-Report, Observer and Neuropsychological Measures
There are no meaningful correlations between the AKGT total score and any of the neuropsychological, self-report, and observer measures in both groups (Table 6). The same holds true for the CAARS-L: S Inconsistency Index.
Correlations (Spearman-rho) Between Validity Measures and Self-Report, Observer, and Neuropsychological Measures.
Note. Correlations for those who failed the AKGT (n = 63), and for those who passed the AKGT (n = 133) in parenthesis. Correlations printed in bold are significant after Bonferroni correction in each group (p = .05 / 145 = .00035). Significant differences (two-sided) between correlations are listed under the respective pairs. AKGT = Amsterdam Short Term Memory Test; CAARS = Conners Adult ADHD Rating Scales; DSM-IV = Diagnostic and Statistical Manual of Mental Disorders (4th ed.; American Psychiatric Association, 1994).
The CAARS-L: O Inconsistency Index has some substantial correlations (−.30 to −.38) in the AKGT fail group, mostly with patient self-report measures. They do not reach significance, are negative, and therefore contrary to expectation. High correlations occur between the CAARS-L: O Infrequency Index and almost all CAARS-L: O subscales in both groups except for “Self-Concept.” Almost the same pattern emerges for the CAARS-L: S Infrequency Index, which also exhibits high associations with the ADHS-SB in both groups. The correlation between the CAARS-L: S subscale “Self-Concept” and the CAARS-L: S Infrequency Index is significantly higher (p = .01) in the AKGT fail group (r = .62) compared with the AKGT pass group (r = .31).
Discussion
In the present study, we analyzed data from patients diagnosed with ADHD who, in addition to the validity scales of the CAARS-L: S/O, also completed a symptom validity test as part of our standard diagnostic strategy. In our sample, 32.1% scored below the cutoff level of the AKGT, which might be an indicator for non-credible performance. Therefore, we compared those who failed this test with those who passed it on self-report, observer, and neuropsychological measures. The two groups did not differ significantly with respect to ADHD self-reported symptoms. Both groups exhibited a high symptom load; the CAARS validity indices were elevated in 15.9% (AKGT fail) and 23.3% (AKGT pass; Inconsistency Index), and in 41.3% (AKGT fail) and 33.1% (AKGT pass; Infrequency Index) in both groups (p > .05). The two groups also did not differ significantly regarding observer ratings of ADHD symptoms and their respective validity indices. The scales of the CAARS observer version correlated between r = .21 and r = .69 with their respective scales of the self-report form in both groups with higher correlations in the AKGT fail group in the subscales “Inattention” and “DSM-IV Inattentive Symptoms.” Differences were found in neuropsychological measures. Those who failed the AKGT had higher reaction time variabilities in selective attention, higher omission errors in selective attention, higher reaction time variabilities in Auditory and Visual divided attention, higher omission errors in sustained attention, and a more atypical result in the Qb+© subdomain Inattention. Compared with population norms, they performed below average in the selective attention variability, the auditory divided attention variability and omission errors, and in sustained attention omission errors.
Intercorrelations between validity measures were mostly not significant and not different between the two groups. There were no meaningful correlations between the AKGT total score and the CAARS-L: S Inconsistency Index and any of the neuropsychological, self-report, and observer measures in both groups. We found high correlations between the CAARS-L: O Infrequency Index and almost all CAARS-L: O subscales in both groups, except for “Self-Concept.” This was also shown in the CAARS-L: S Infrequency Index, which also exhibited high associations with the ADHD Self-Rating Scale in both groups.
The failure rate in our study in a symptom validity measure with 32.1% closely corresponds to those of other studies: 31% (Suhr et al., 2008) and 22% (Marshall et al., 2010). Our sample differs from these studies in that our patients are older. One interpretation might be that this proportion represents a base rate of negative response bias in individuals seeking evaluation for ADHD. However, this might be a subgroup with profound neuropsychological deficits, for example, fluctuations in attention, who have difficulties with even basic tests.
The majority of the studies did not find differences in self-report measures between instructed malingerers and control groups or between those who failed the symptom validity tests and those who passed them (Musso & Gouvier, 2014; Tucha et al., 2014). We also found no differences in self-report measures and CAARS validity indices between our ADHD patients who failed the AKGT and those who passed the test. This might be regarded as an ability of a non-credible group to feign symptoms of ADHD that were highly manifested in both groups. However, one may argue that ADHD patients scoring below the cutoff of a symptom validity test do not seem to exaggerate their symptomatology. Failure on a symptom validity test does not automatically indicate that an individual is showing a negative response bias (J. Suhr et al., 2008). It is indeed problematic that a considerable amount of students instructed to feign ADHD were not identified by symptom validity tests (Jasinski et al., 2011; Sollman, Ranseen, & Berry, 2010). Suhr et al. (2008) found that the CAARS Inconsistency Index was not specific for their non-credible group. The sensitivity for detecting malingered ADHD was low (Musso & Gouvier, 2014). In our sample, only 11 (5.6%) patients scored higher than both cutoff values of the Inconsistency and the Infrequency Indices. This further demonstrates the limited usefulness of these indices, or it shows that the majority responds in a stringent way. To further complicate things, there are no special validated strategies to detect sub-optimal effort in ADHD (Lee Booksh et al., 2010).
To our knowledge, this study is the first using observer ratings in the context of symptom validity testing in adults with ADHD. We found no significant differences in observer ratings and CAARS validity indices between the two groups with high versus low AKGT scores. This finding might be an indicator for credible performance in both groups, although one has to consider that a closely associated person might rate in favor of a friend or relative seeking evaluation for ADHD.
The higher reaction time variabilities in Selective, Auditory and Visual divided attention, and higher omission errors in sustained attention in those who failed the AKGT might represent a special attentional deficit in this group. This is further supported by the below-average performance in the Auditory divided attention variability and omission errors, and in Sustained attention omission errors. The underperformance in Auditory divided attention reveals a special deficit in divided attention in this group. Increased reaction time variability might be an indicator for temporal processing deficits, general problems in maintaining alertness, and focused attention, and it might also be a suitable endophenotype for ADHD (Adams et al., 2011; Kofler et al., 2013; Uebel et al., 2010). There is currently no neuropsychological profile that is characteristic for ADHD (Musso & Gouvier, 2014).
Intercorrelations between validity measures were mostly insignificant, advocating for the argument that the patients did not continuously show negative response bias. Contrary to this, a considerable amount in both samples scored above the cutoff values of the CAARS Inconsistency and Infrequency Indices. These indices and the total score of the AKGT did not significantly correlate with the neuropsychological measures, making it rather improbable that the patients performed intentionally worse in these examinations. What warrants further attention are the high correlations between CAARS-L: S and O subscales and their respective Infrequency Index in both groups. This index is still in the developmental phase; therefore, it is too early to conclude whether this is a sign of non-credible performance (Suhr, Buelow, et al., 2011).
To conclude, we found no strong indicators for a negative response bias in those ADHD patients who failed a symptom validity test and those who passed, but one must consider that a proportion of those passing these tests might nevertheless exhibit such behavior.
Strengths and Limitations
Only a few studies examined non-credible performance in clinical samples of ADHD patients. Therefore, a strength of our study is the analysis of a sample of ADHD patients. We performed a thorough assessment of ADHD with multiple strategies, such as a structured clinical interview, self-report and observer rating scales, and neuropsychological tasks, including a CPT (Quinn, 2003).
It is recommended to apply several symptom validity tests. Musso and Gouvier (2014) stated that failure of three or more symptom validity tests proved most useful at detecting malingered ADHD. In our study, we used only one such test because we had to limit the examination time. The Inconsistency and Infrequency Indices of the CAARS are explicitly promoted as validity scales, but we neither found significant differences between groups according to those measures nor significant correlations between those indices and the AKGT.
Our analyses demonstrate that the detection of negative response bias in adult ADHD patients is a complicated matter as symptom validity tests and validity indices of self-report measures have significant limitations in uncovering such behavior. Consequently, and in agreement with Tucha et al. (2014), new measures and approaches to detect feigned ADHD should be developed.
Footnotes
Authors’ Note
The data were presented as a poster at the 3rd EUNETHYDIS International Conference on ADHD, May 21 to 24, 2014, in Istanbul, Turkey.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
