Abstract
The current study examined the incremental utility of rating scales, a structured diagnostic interview, and multiple informants in a comprehensive assessment of attention-deficit/hyperactivity disorder (ADHD). The sample included 185 children with ADHD (Mage = 9.22, SD = 0.95) and 82 children without ADHD (Mage = 9.24, SD = 0.88). Logistic regressions were used to examine the incremental contribution of each method within an assessment of ADHD. Results indicated that information collected from a structured diagnostic interview was unable to significantly improve a prediction model including parent and teacher ratings (Block χ2 = 0.91, ρ = .64). Teacher ratings on symptom-based scales resulted in significant model improvement beyond parent ratings alone (Block χ2 = 48.47, ρ < .001). Exploratory analyses indicated that using behavioral rating scales correctly classified all participants by diagnosis. Clinical implications are highlighted, and future research directions are discussed.
Attention-deficit/hyperactivity disorder (ADHD) is a disorder characterized by a persistent pattern of developmentally inappropriate levels of inattention, hyperactivity, and impulsivity (American Psychiatric Association [APA], 2000). Current Diagnostic and Statistical Manual of Mental Disorders (4th ed., text rev.; DSM-IV-TR; APA, 2000) requirements include an assessment of functioning in multiple settings, documented age of onset and symptom criteria, and associated impairment to be eligible for a diagnosis of ADHD. Past research has primarily focused on identifying the core symptoms in multiple settings for the purpose of establishing a diagnosis of ADHD through the use of a multimethod assessment battery that may include parent and child diagnostic interviews, behavioral rating scales completed by parents and/or teachers, direct observations, and/or clinic-based assessments (e.g., continuous performance tasks; Pelham, Fabiano, & Massetti, 2005). Although research strongly supports the integration of information across settings and informants (Pelham et al., 2005; Power, Doherty, et al., 1998; Tripp & Clarke, 2006), fulfilling these recommendations leads to an emphasis on a multi-informant, multimethod assessment that requires extensive cost, time, and resources (Johnston & Murray, 2003). In fact, despite an abundance of research documenting the reliability and validity of numerous rating scales and structured interviews and calls for more efficient, cost-effective assessments of ADHD, little, if any, research has examined the actual incremental validity and clinical utility of these methods within a comprehensive assessment of ADHD (Johnston & Murray, 2003; Pelham et al., 2005; Wright, Waschbusch, & Frankland, 2007).
The current study addresses this gap in the literature through the examination of the incremental and clinical utility of assessment methods demonstrating the most promising evidence of efficiency and predictive validity: symptom-based rating scales, empirically derived rating scales, and a structured diagnostic interview.
Current Assessment Strategies for ADHD
As noted, the pervasive and chronic nature of problems associated with ADHD requires an evaluation that involves collecting data across multiple settings and caregivers.
Though originally designed for epidemiological studies, the use of structured diagnostic interviews in the clinical realm has become increasingly common due to their ability to assess a wide range of DSM-IV-TR disorders (Diagnostic and Statistical Manual of Mental Disorder 4th ed., text rev.; DSM-IV-TR; APA, 2000) in a consistent, standardized manner (Hodges, 1993). This structure was designed to reduce clinicians’ tendencies to collect information selectively (leading to an incorrect diagnosis), make diagnoses most familiar to them, or determine a diagnosis before all relevant information is collected (McClellan & Werry, 2000). Clinicians have used structured interviews as an indicator of initial diagnosis and treatment response, and as a comparison for prior diagnoses using other strategies (Piacentini et al., 1999). However, researchers have noted significant decreases in symptom endorsement over the course of administration (attenuation effects), influence of interviewer and participant characteristics on diagnosis, and low to moderate levels of test–retest reliability (Jensen & Edelbrock, 1999; Piacentini et al., 1999; Roberts, Solovitz, Chen, & Casat, 1996). Furthermore, structured diagnostic interviews require considerable time to administer, score, and interpret. For example, the Diagnostic Interview Schedule for Children-IV (DISC-IV) is composed of 358 “stem questions” and almost 3,000 additional questions that may be administered (Shaffer, Fisher, Lucas, Dulcan, & Schwab-Stone, 2000).
Conversely, the ease and efficiency of numerous symptom-based and empirically derived rating scales has led to their frequent use in the assessment process (Pelham et al., 2005). Rating scales allow clinicians to gather information efficiently (typically less than 20 min) from informants who have known the child for months or years (Power, Doherty, et al., 1998). In particular, multiple DSM-IV symptom-based rating scales have been shown to provide information about behavior in multiple settings (e.g., home and school), to discriminate between clinical and nonclinical groups and between subtypes of ADHD, and to be sensitive to both behavioral and pharmacological treatment effects (for a review, see MTA Cooperative Group, 1999; Pelham et al., 2005; Power, Andrews, et al., 1998; Power, Doherty, et al., 1998). These measures typically use each of the 18 symptoms defining ADHD on a Likert-type scale and include norms and/or suggested cut points for a diagnosis of ADHD. Empirically derived rating scales also have been shown to accurately identify children with ADHD, discriminate children with inattention only from children with inattention and hyperactivity, and be sensitive to treatment effects (for a review, see Pelham et al., 2005). Of particular interest to this study, the Attention Problems Syndrome Scales on the Child Behavior Checklist (CBCL) and Teacher Report Form (TRF) have been used as a proxy for diagnosis in several studies, are highly related to DSM-IV ADHD diagnoses, are able to discriminate between subtypes of ADHD, and have been used to assess outcomes in numerous treatment studies (Chen, Faraone, Biederman, & Tsuang, 1994; Hartman, Stage, & Webster-Stratton, 2003).
Research has shown that combining parent and teacher ratings of ADHD symptoms increases the sensitivity and specificity of a measure beyond that provided by one informant (Loeber et al., 1991; Power, Doherty, et al., 1998; Tripp & Clarke, 2006). However, other studies have reported low rates of specificity, high rates of cross-informant discrepancies (e.g., raters from different settings), and low rates of interobserver agreement (e.g., raters from the same setting) in evaluations using behavioral rating scales (Collett, Ohan, & Myers, 2003; Reid & Maag, 1994; Wender, 2004).
Although these findings provide support for the reliability and validity of structured diagnostic interviews, symptom-based and empirically derived rating scales, and multiple informants, sparse research has examined their incremental contribution within an assessment of ADHD. Indeed, Pelham and colleagues (2005) argued that structured diagnostic interviews provide little information beyond that gathered through more efficient methods (e.g., behavioral rating scales) in the assessment of ADHD. In their review, they cite research noting the efficiency and utility of shorter empirically derived scales, and/or subsets or even individual items in identifying children with ADHD (Pelham et al., 2005). Furthermore, they review studies documenting high correlations between structured diagnostic interviews and rating scales and near perfect rates of agreement when classifying children using rating scales or a structured diagnostic interview (DuPaul, Power, McGoey, Ikeda, & Anastopoulos, 1998; Ostrander, Weinfurt, Yarnold, & August, 1998; Power, Costigan, Leff, Eiraldi, & Landau, 2001). Pelham and colleagues go on to argue that “diagnosing ADHD is most efficiently accomplished with parent and teacher rating scales” (Pelham et al., 2005, p. 469). Other researchers have noted the surprising lack of research examining the unique or additional information provided by different methods, measures, or informants arguing that clinicians tend to have “blind faith in the ‘more is better’ approach” (Johnston & Murray, 2003, p. 500). However, only one study has examined the actual agreement between ratings of ADHD symptomatology on behavioral ratings scales and a structured diagnostic interview (Wright et al., 2007). However, their study examined neither teacher ratings nor the incremental utility of rating scales within a diagnostic formulation.
One study examining assessments using parent- and teacher-rated symptoms of ADHD concluded that the optimal approach varied as a function of informant, scale, and purpose of assessment (Power et al., 2001). However, although their study included a structured diagnostic interview, they did not examine the incremental utility of parent and teacher report in conjunction with the structured diagnostic interview. Hence, despite the wealth of evidence supporting the differing predictive power of ADHD symptoms by method and by informant (Power, Andrews, et al., 1998; Power, Doherty, et al., 1998; Wolraich et al., 2004), few studies have examined the incremental contribution of different types of scales, structured diagnostic interviews, and/or informants thereby failing to answer fully this particular question of clinical utility and efficiency. Past research has either (a) compared the utility of different raters in predicting a diagnosis of ADHD or (b) examined the correlations among behavioral rating scales and structured diagnostic interviews. Research must examine the relative incremental validity of a measure in relation to its additive contribution beyond that which may be predicted by other, more cost-effective measures or methods. Thus, analyses examining the degree of correlation among multiple methods and raters as well as analyses examining the actual, independent contribution each adds in the assessment of ADHD are needed. Furthermore, each method also must demonstrate clinical utility in real-world settings to justify its efficiency and cost-effectiveness. Only studies using this type of design are able to inform the relation between incremental utility and cost-effectiveness (Johnston & Murray, 2003).
Study Rationale and Hypotheses
The goals of this study were twofold: (a) to examine the correlation and rate of agreement among different informants and methods in relation to ADHD symptoms and (b) to examine the incremental utility of symptom-based rating scales, empirically derived rating scales, and a structured diagnostic interview within a multimethod, multi-informant assessment of ADHD. Exploratory analyses also were conducted to examine the relative incremental utility and efficiency of clinically relevant algorithms using the methods demonstrating the greatest association with a diagnosis of ADHD.
Given reported low rates of specificity and interobserver agreement and high rates of cross-informant discrepancies, correlations and agreement between ratings of ADHD symptoms on behavioral rating scales and a structured interview and a consensus diagnosis derived using information culled from multiple sources/methods were examined. Consistent with previous research and study design, we hypothesized that each method would be independently associated with a diagnosis of ADHD (Chen et al., 1994; Power, Andrews, et al., 1998; Power, Doherty, et al., 1998; Schaffer et al., 2000).
Second, we examined the relative incremental contributions of each reviewed method in the prediction of a diagnosis of ADHD. Given the evidence supporting the increased predictive validity of models including multiple informants (e.g., parent and teacher; Power, Doherty, et al., 1998; Power et al., 2001), we hypothesized that teacher ratings would significantly improve predictive models using parent ratings alone. We also posited that a structured diagnostic interview would not account for significant variance beyond that accounted for by the more efficient parent and teacher rating scales providing empirical support for Pelham and colleagues’ (2005) conclusion that these methods provide redundant information. Furthermore, given the redundancy in informants when using parent-completed rating scales and a structured diagnostic interview (also completed by the parent), we hypothesized that a structured diagnostic interview would not significantly improve a model including parent-completed rating scales. Alternatively, we hypothesized that a structured diagnostic interview would significantly improve a model including teacher-completed rating scales given the nonredundancy in informants. Despite near universal endorsement for their inclusion in the assessment of ADHD, little empirical evidence exists to support the inclusion of structured diagnostic interviews within an assessment of ADHD that also includes behavioral rating scales. As such, to our knowledge, this is the first study to examine the incremental utility of a structured diagnostic interview beyond that of the more efficient rating scales. Of note, this study examined only the ADHD portion of the DISC interview. The validity of the DISC interview in discriminating and/or diagnosing other disorders was not examined in the present study.
We also wished to explore the utility and efficiency of clinically relevant diagnostic algorithms based on the above statistical models in classifying children when compared with a diagnosis of ADHD derived using a comprehensive “gold standard” assessment. As argued by Johnston and Murray (2003), research regarding incremental validity “must be conducted with procedures, measures, and samples that reflect the realities of clinical practice” (p. 504). Given the lack of expediency in using statistical models in a clinic setting, diagnostic algorithms may provide practical, real-world procedures to use most efficiently the various methods informing the assessment of ADHD.
Given current controversy in the literature regarding whether ADHD–predominantly inattentive type (ADHD-I) should be considered a subtype of ADHD given the unique etiology, core deficits, associated features, and comorbid functioning of children diagnosed with ADHD-I (Milich, Balentine, & Lynam, 2001), only children with either ADHD–combined type (ADHD-C) or ADHD–predominantly hyperactive/impulsive type (ADHD-HI) were examined in this study. Given this selection bias, we expected that ratings of hyperactivity/impulsivity would account for greater variance than ratings of inattention in predicted models.
Method
Participants
Participants were drawn from an existing database from a study funded by the National Institute of Mental Health addressing unrelated research questions. Data collection occurred at three different sites including a large, public midwestern university and two large- to midsized, public northeastern universities. Informed consent was obtained for all participating families using procedures approved by local Institutional Review Boards at each site. Participants were children aged 7 to 11 years of age. Recruitment was designed to gain a representative sample through the use of school settings, primary medical care settings, mental health practitioners, and self-referrals solicited through advertisements and word of mouth. Participants who met criteria for ADHD-C or ADHD-HI following the assessment procedures discussed below were placed in the ADHD group (n = 185). Children not meeting criteria for ADHD were placed in the control group (n = 82). Children with a Brief Intellectual Ability standard score below 80 on the Woodcock-Johnson Test of Cognitive Abilities (WJ-TC; Woodcock, McGrew, & Mather, 2001), with a previous diagnosis of any pervasive developmental disorder, or who were currently taking medications that affected behavior and that could not be withdrawn for testing were excluded from the study. Children meeting criteria for ADHD-I also were ineligible for the study.
Consensus diagnosis
All participants received a comprehensive assessment of ADHD that followed current “gold standard” guidelines as established by the American Academy of Pediatrics (2000). Information was gathered through the use of a semistructured clinical interview (regarding developmental, social, academic, and family functioning), parent and teacher symptom-based and empirically derived rating scales, and a comprehensive structured diagnostic interview. Each child received an assessment battery including (a) a cognitive and achievement battery; (b) self-report measures of self-perception, anxiety symptoms, and depressive symptoms; and (c) a clinical interview. Of note, the structured diagnostic interview was administered to each parent following completion of the behavior rating scales and semistructured interview with the ADHD module administered at the end of the interview as designed. Assessments typically were completed during one or two sessions by the child and his or her parent. Teacher-completed ratings were mailed to each participant’s teachers following informed consent procedures with the parent. Clinicians administering the assessment were PhD-level psychologists or trained graduate-level research assistants.
Diagnostic decisions were made by five PhD-level clinical or school psychologists who worked in pairs and who specialized in externalizing disorders in childhood (hereafter referred to as the consensus diagnosticians). Four of the five consensus diagnosticians were blind to the purpose of the present study, whereas one consensus diagnostician (the second author) was not. Each participant’s assessment data were reviewed by two consensus diagnosticians who derived a diagnosis for each child independently. The assessment data reviewed included the semistructured clinical interview, parent- and teacher-completed Disruptive Behavior Disorders (DBD) Rating Scales (Pelham, Gnagy, Greenslade, & Milich, 1992), the Computerized DISC-IV parent version (Shaffer et al., 2000), the CBCL and TRF (Achenbach & Rescorla, 2001), and the Woodcock-Johnson III Tests of Cognitive Abilities (WJ-TC) and Achievement (WJ-TA; Woodcock et al., 2001). The Test Observation Form (TOF; McConaughy & Achenbach, 2004) was also available to the clinicians for a subset of children (n = 86) at one site. When diagnostic decisions were not in agreement, the two consensus diagnosticians who had reviewed the child’s file discussed the participant’s assessment data until a consensus diagnosis was agreed on. If a consensus regarding diagnosis could not be reached, the child was dropped from the study and all analyses (n = 1). Consensus diagnosticians reported using parent- and teacher-completed ratings and the structured diagnostic interview relatively equally with less reliance on measures of cognitive ability and achievement in relation to a diagnosis of ADHD. Furthermore, DSM-IV criteria were followed, including symptom requirements, impairment in two or more settings, and onset prior to age 7. As discussed below, the consensus diagnosis was used as the criterion in the present study (e.g., ADHD or not ADHD).
Parent Measures
Demographic Questionnaire (DQ)
Demographic information including parental income, educational level, occupation, and marital status was provided by each participant’s caretaker.
DBD Rating Scale
The DSM-IV (Massetti et al., 2003: Pelham et al., 1992) version of the parent DBD Rating Scale is a measure of the DSM-IV-TR (APA, 2000) symptoms of ADHD, Oppositional Defiant Disorder (ODD), and Conduct Disorder (CD). The measure also includes some Diagnostic and Statistical Manual of Mental Disorders (3rd ed., rev.; DSM-III-R; APA, 1987) symptoms of ADHD. Each of the 45 items is rated by the parent on a 4-point scale ranging from 0 (not at all present) to 3 (very much present). Items rated as a 2 (pretty much present) or 3 (very much present) are considered to be endorsed as symptoms. For the present study, the nine inattention and nine hyperactivity/impulsivity symptoms from the DSM-IV version of the DBD were used (comprising the Inattention and Hyperactivity/Impulsivity Scales, respectively). The coefficient alphas for the present sample were .93 for the Inattention Scale and .96 for the Hyperactivity/Impulsivity Scale.
CBCL
The CBCL (Achenbach & Rescorla, 2001) is a 118-item parent-completed behavior problem checklist using a Likert-type scale designed to assess multiple domains of children’s externalizing and internalizing functioning in the past 6 months. The CBCL has been used widely for obtaining ratings of problem behavior in children and has demonstrated strong evidence of reliability and validity (Achenbach & Rescorla, 2001). Syndrome scales on the CBCL were derived through factor analyses. Only the Attention Problems Syndrome Scale was used in this study. The coefficient alpha for the Attention Problems Syndrome Scale for the present sample was .89.
The Computerized DISC-IV
The DISC-IV (Shaffer et al., 2000) is a structured diagnostic interview designed for use by lay interviewers in epidemiological studies to elicit DSM-IV-TR and the International Classification of Diseases–Tenth Revision diagnoses for children and adolescents covering 36 mental health disorders (Shaffer et al., 2000). There are 358 “stem” questions that are asked of every respondent who are overly sensitive to lead to more “contingent” questions that are able to differentiate true positives from false positives. This study used the ADHD module in the computerized version of the DISC-IV. Research has supported the concurrent criterion validity of the ADHD module in relation to other diagnostic interviews, symptom checklists, and external validators (e.g., school dysfunction, functional impairment; Cohen, O’Conner, Lewis, Velez, & Malachowski, 1987; Jensen et al., 1996; Piacentini et al., 1993). For a symptom to be endorsed as present, criterion questions examine the relative frequency and duration of the symptom in multiple settings. The number of inattention and hyperactivity/impulsivity symptoms identified as present on the DISC-IV was examined in this study. The coefficient alphas for the present sample were .93 for the Inattention Scale and .90 for the Hyperactivity/Impulsivity Scale.
Teacher Measures
DBD Rating Scale
The teacher DBD Rating Scale (Massetti et al., 2003; Pelham et al., 1992) was identical to the parent DBD Rating Scale. The Inattention and Hyperactivity/Impulsivity subscales were used from the teacher version. The coefficient alphas for the present sample were .95 and .95 for the Inattention and Hyperactivity/Impulsivity Scales, respectively.
TRF
The TRF (Achenbach & Rescorla, 2001) is a 118-item teacher-completed behavior problem checklist using a Likert-type scale to assess multiple domains of children’s internalizing and externalizing functioning in the past 6 months. The TRF has strong evidence of reliability and validity and has been widely used for obtaining ratings of problem behavior in children (Achenbach & Rescorla, 2001). For this study, only the empirically derived narrow-band Attention Problems Syndrome Scale was used. The coefficient alpha for the Attention Problems Syndrome Scale in the present sample was .87.
Statistical Analyses
Univariate ANOVA and chi-square tests were conducted examining group differences across a range of demographic variables. As the study measures (DBD, CBCL, TRF, DISC) were also used by the consensus diagnosticians in deriving the criterion (i.e., diagnosis), substantial bias exists in relation to tests of predictive validity. However, as the consensus diagnosticians had access to all measures and reported using measures relatively equally, this bias may be assumed equivalent across measures. Nevertheless, correlational analyses and measures of agreement (κ) were conducted to examine relations among ADHD symptoms as endorsed on parent- and teacher-completed DBD Rating Scales and the parent-completed DISC interview to assess collinearity among predictors.
A series of hierarchical logistic regression analyses were conducted examining the relative incremental utility of each method and informant in contributing to the consensus diagnosis. This approach examines the relative incremental utility of each assessment method within an assessment for ADHD (in relation to the other methods used within the same assessment). For each of the data analytic procedures used, mean scores were calculated for parent and teacher ratings on the DBD Inattention and Hyperactivity/Impulsivity Scales (PDBD-IA, PDBD-HI; TDBD-IA, TDBD-HI), T-scores were used from the CBCL and TRF Attention Problems Scales (CBCL-A, TRF-A), and the total number of symptoms of inattention and hyperactivity/impulsivity endorsed as present on the DISC (DISC-IA, DISC-HI) were computed. Logistic regressions examined the relative improvement in diagnostic accuracy following the addition of multiple methods and/or informants. The Likelihood Ratio Chi-Square test was used to test the significance of each additional block/step (Block χ2). The Hosmer and Lemeshow Chi-Square test (Hosmer & Lemeshow, 2000) and Nagelkerke’s R2 were examined to ensure adequate model fit and relative strength of association between predictor and criterion variables, respectively.
Finally, a series of exploratory analyses was conducted examining the clinical utility of diagnostic algorithms using the measures that demonstrated the greatest association with consensus diagnoses in the above logistic regression models. Algorithms were examined in order of efficiency (i.e., most efficient to least efficient), until classification and sensitivity/specificity rates were equivalent to the classification rates of the above logistic regression models. Sensitivity, specificity, percentage agreement, and chance-corrected agreement were calculated for each algorithm.
Results
Preliminary Analyses
Preliminary analyses were conducted comparing groups (e.g., ADHD or not ADHD) with regard to sex, a diagnosis of ODD/CD, grade retention, race/ethnicity, cognitive ability, Internalizing and Externalizing Problems scores from the CBCL and TRF, and age. Categorical variables were compared using chi-square analyses, and univariate ANOVAs were conducted for continuous variables. Chi-square analyses indicated that the children in the ADHD group were more likely to be male, had significantly higher rates of ODD or CD diagnoses, and were significantly more likely to have been retained than children in the control group; no group differences were found related to race/ethnicity (Table 1). Univariate ANOVA analyses indicated that the ADHD group had significantly lower standard scores on the Brief Intellectual Ability Scale of the WJ-TC and significantly higher rates of Internalizing and Externalizing Problems on both the CBCL and TRF (as indexed by T-scores) compared with the control group (Table 1).
Summary of Univariate Analysis of Variances Comparing Children With and Without ADHD on Measures of IQ, Internalizing and Externalizing Problems, and Age.
Note. ADHD = attention-deficit/hyperactivity disorder; ODD/CD = oppositional defiant disorder/conduct disorder; BIA = Brief Intellectual Ability standard score as indexed on the Woodcock-Johnson III Tests of Cognitive Abilities; CBCL-Int = Child Behavior Checklist–Internalizing Problems Broadband Scale; CBCL-Ext = Child Behavior Checklist–Externalizing Problems Broadband Scale; TRF-Int = Teacher Report Form–Internalizing Problems Broadband Scale; TRF-Ext = Teacher Report Form–Externalizing Problems Broadband Scale.
Using the assessment measures, diagnoses of ADHD were derived using ratings on the parent and teacher DBD Rating Scale and/or diagnoses received on the parent-completed DISC interview. Agreement between the DBD and DISC-derived diagnoses and the consensus diagnosis was examined. As the DISC also was used to derive consensus diagnoses, this analysis was completed to examine whether diagnosticians followed DISC-derived diagnoses solely in formulating consensus diagnoses (100% agreement between DISC and consensus diagnosticians’ diagnoses would contraindicate the validity of incremental utility analyses). As expected, diagnostic agreement was found for the majority of cases (89.1%; κ = .768). However, clinicians disagreed with the DISC diagnosis (or lack thereof) for more than 10% of cases. In comparison, using PDBD and requiring a minimum of six inattentive and/or six hyperactive/impulsive symptoms to be classified as ADHD, agreement was found for 86.5% of cases (κ = .718). Thus, each method provided some unique information in deriving consensus diagnoses.
Primary Analyses
First-order correlations and symptom agreement were examined between the symptoms endorsed on the DBD Rating Scales and the DISC interview. Correlations between all parent-rated ADHD symptoms on the DBD and parent report on the DISC were significant (Table 2). Although less highly correlated, teacher-rated symptoms on the DBD were significantly associated with parent report on the DISC as well (Table 2).
Summary of Correlations, Percentage Agreement, and Chance-Corrected Agreement (κ) Between Parent and Teacher Ratings of ADHD Symptoms on the Disruptive Behavior Disorders and Parent Report on the Diagnostic Interview Schedule for Children-IV.
Note. ADHD = attention-deficit/hyperactivity disorder.
p ≤ .05. **p ≤ .01. ***p ≤ .001.
Actual and chance-corrected (κ) agreement was calculated for each symptom between the DBD Rating Scale and DISC interview. Following the findings of Power and colleagues (Power, Andrews, et al., 1998; Power, Doherty, et al., 1998; Power et al., 2001) and general consensus (Pelham et al., 2005), symptoms endorsed as “pretty much” or “very much” on the DBD Scales were considered present. Chance-corrected symptom agreement between PDBD and the DISC interview ranged from moderate to large across both symptoms of inattention and hyperactivity/impulsivity (Table 2). Whereas, chance-corrected agreement between TDBD and DISC ranged from small to moderate (Table 2).
Agreement between the DBD and DISC-derived diagnoses also was examined. Specifically, DBD diagnoses were derived following DSM-IV criteria by requiring a minimum of six symptoms on either of the ADHD dimensions to be endorsed as present to be classified as ADHD. Agreement was examined between DBD-derived diagnoses and DISC diagnoses. Agreement was found between PDBD and report on the DISC for 90.6% of cases (κ = .809), whereas 76.8% (κ = .535) of participants received the same diagnosis (or lack thereof) when TDBD and parent report on the DISC were examined. When parent and teacher ratings were used together (a minimum of six symptoms endorsed by either rater in either dimension), PDBD and TDBD, agreed with the DISC diagnosis on 89.9% of cases (κ = .786). Lower agreement was due to more cases classified as ADHD by the combined DBD ratings.
Incremental Utility of Symptom-Based and Empirically Derived Rating Scalesby Informant
Logistic regressions examined the contribution of empirically derived rating scales beyond that of symptom-based rating scales. The more efficient symptom-based rating scales were entered into the model first. Results indicated that the addition of the CBCL-A to a model including PDBD resulted in statistically significant model improvement, Block χ2 (2, N = 267) = 6.37, ρ = .012. In contrast, the addition of the TRF-A to a model including TDBD did not result in significant model improvement, Block χ2 (1, N = 267) = 0.76, ρ = .383. All regression models adequately fit the data.
Incremental Utility of Parent- and Teacher-Completed Rating Scales
The second set of logistic regressions examined the incremental utility of PDBD and TDBD in the prediction of consensus diagnosis. As assessments of ADHD almost uniformly include parent ratings of a child’s behavior, we examined the incremental utility of teacher ratings beyond that of parent ratings. Furthermore, as symptom-based rating scales are more efficient than the longer, empirically derived scales, they were entered into the model first.
PDBD-IA and PDBD-HI were entered followed by TDBD-IA and TDBD-HI at Step 2. At Step 3, we included the CBCL-A and TRF-A. Results indicated that the addition of TDBD did result in statistically significant model improvement, Block χ2 (2, N = 267) = 48.47, ρ < .001. The addition of ratings on the CBCL-A and TRF-A did not significantly improve the overall model, Block χ2 (2, N = 267) = 0.59, ρ = .756. All regression models adequately fit the data.
Given the robustness of the PDBD and TDBD in classifying cases, a post hoc conditional forward-entry logistic regression analysis was conducted including each of the parent and teacher DBD Scales to examine their respective statistical contribution in the prediction of consensus diagnosis. Results indicated that PDBD-IA entered the model first, followed by TDBD-HI at Step 2 and PDBD-HI at Step 3. TDBD-IA were not incrementally associated with consensus diagnosis.
Incremental Utility of a Structured Diagnostic Interview
The final set of logistic regression analyses examined the incremental utility of a structured diagnostic interview and parent- and/or teacher-completed rating scales. The DISC Inattention Scale (DISC-IA) and DISC Hyperactivity/Impulsivity Scale (DISC-HI) significantly improved a model including the PDBD Scales alone, Block χ2 (2, N = 267) = 9.85, ρ = .007. The addition of the DISC-IA and DISC-HI to teacher-completed rating scales also indicated statistically significant model improvement, Block χ2 (2, N = 267) = 98.97, ρ < .001. As hypothesized, the addition of the DISC-IA and DISC-HI Scales within a model including parent- and teacher-completed symptom-based rating scales neither resulted in significant model improvement, Blockχ2 (2, N = 267) = 0.91, ρ = .636, nor increased the rate of classification. In contrast, when the DISC-IA and DISC-HI were entered into the model first, the addition of parent- and teacher-completed symptom-based rating scalesdid result in significant model improvement, Blockχ2 (7, N = 267) = 61.761, ρ < .001.
It should be noted that covariates were not included in the regression models despite significant group differences (namely, sex or IQ). Past reviews suggest that analysis of covariance is inappropriate when applied to situations involving nonrandom group assignments, as preexisting group differences cannot then be assumed to be independent of the predictor variables (Miller & Chapman, 2001). For the skeptical reader, however, all regression models also were conducted including sex and IQ as covariates. These models did not differ markedly from the above regression models and, as such, are not reported here.
As predictor variables were also used in deriving the criterion (i.e., consensus diagnosis), regression models resulted in extremely large odds ratios and insignificant Wald statistics (Albert & Anderson, 1984; Heinz & Schemper, 2002). Resulting odds ratios and Wald statistics were neither interpretable nor germane to the question of relative incremental utility addressed here. For these reasons, odds ratios and Wald statistics are not reported here.
Examination of the Clinical Utility of Diagnostic Algorithms
Diagnostic algorithms followed procedures outlined above (e.g., endorsement of a minimum of six symptoms on each of the ADHD dimensions). We initially required that the algorithm derive diagnostic classifications parallel with the consensus diagnoses for agreement (e.g., endorsement of at least six symptoms of both inattention and hyperactivity/impulsivity). We then examined the rate of agreement between consensus diagnosis and any algorithm-derived diagnosis of ADHD (e.g., inattentive type only). The second “lenient” algorithm required the endorsement of six symptoms of inattention or hyperactivity/impulsivity to receive a classification as ADHD. Chance-corrected rates of agreement were then calculated for the “stringent” and “lenient” algorithm-derived classifications with consensus diagnosis. Algorithms using multiple informants followed a flexible algorithmic approach similar to that used in previous studies of ADHD (Rowland et al., 2001; Wolraich et al., 2004). To ensure the presence of symptoms in multiple settings, we required a minimum of six symptoms to be endorsed as present by one informant with at least three symptoms endorsed by the other informant.
Initial algorithms used the most efficient methods/measures that also demonstrated incremental utility in the logistic regression models (e.g., parent-completed symptom-based rating scales alone). We then proceeded to use additional measures in subsequent algorithms. Thus, we continued deriving less-efficient, but increasingly comprehensive algorithms until maximal agreement between the consensus and algorithm-derived diagnoses was met. As seen in Table 3, algorithms using “lenient” criteria consistently displayed higher rates of sensitivity and chance-corrected agreement. Furthermore, the lenient diagnostic algorithm using all behavioral rating scales resulted in perfect agreement with consensus diagnosis.
Summary of Sensitivity, Specificity, Percentage Agreement, and Chance-Corrected Agreement for Diagnostic Algorithms Using Parent and Teacher Ratings on the Disruptive Behavior Disorders Rating Scales, the Child Behavior Checklist, and the Teacher Report Form.
Note. PDBD-IA = Parent Disruptive Behavior Disorders Rating Scale–Inattention subscale; PDBD-HI = Parent Disruptive Behavior Disorders Rating Scale–Hyperactivity/Impulsivity subscale; TDBD-IA = Teacher Disruptive Behavior Disorders Rating Scale–Inattention subscale; TDBD-HI = Teacher Disruptive Behavior Disorders Rating Scale–Hyperactivity/Impulsivity subscale; CBCL-A = Child Behavior Checklist–Attention Problems Syndrome Scale; TRF-A = Teacher Report Form–Attention Problems Syndrome Scale.
Discussion
The goals of this study were twofold: (a) to examine the correlation and rate of agreement among different informants and methods in relation to ADHD symptoms and (b) to examine the incremental utility of symptom-based rating scales, empirically derived rating scales, and a structured diagnostic interview within a multimethod, multi-informant assessment of ADHD. Exploratory analyses also demonstrated the relative incremental utility and efficiency of clinically relevant algorithms using the methods demonstrating the greatest incremental association with a diagnosis of ADHD.
Consistent with our expectations and previous research (Pelham et al., 2005), our results consistently supported the superiority of both parent-completed methods and of symptom-based rating scales when examined incrementally. Results were less consistent regarding the incremental utility of empirically derived rating scales beyond symptom-based ratings. Specifically, the Attention Problems Syndrome Scale on the CBCL was related to improved model fit, whereas ratings on the Attention Problems Syndrome Scale on the TRF did not result in significant model improvement. As hypothesized, the ADHD portion of a structured diagnostic interview did not significantly improve a regression model containing symptom-based rating scales completed by a child’s parent and teacher. In fact, using both parent and teacher ratings of ADHD symptomatology accounted for virtually all variance within the model. This finding provides empirical support for Pelham’s assertion that a structured diagnostic interview accounts for little incremental utility beyond that accounted for by the more efficient parent and teacher symptom-based rating scales within an assessment of ADHD (Pelham et al., 2005). Of note, the structured diagnostic interview did account for significant improvement in logistic regression models containing either parent or teacher ratings on symptom-based and empirically derived scales. However, as logistic regression models using symptom-based ratings were particularly robust in predicting diagnosis, information provided by a structured diagnostic interview, although statistically significant, accounted for minimal increases in actual diagnostic classification when added to a model including teacher ratings or a model including parent ratings. Importantly, although the structured diagnostic interview was unable to add significant, unique information beyond that provided by parent and teacher ratings, the reverse was not supported. Namely, parent and teacher ratings did account for significant unique information beyond that provided by the structured diagnostic interview.
Following Pelham and colleagues’ assertion that “diagnosing ADHD is most efficiently accomplished with parent and teacher rating scales” (Pelham et al., 2005, p. 469), our findings indicated that using only three nine-item scales (PDBD-IA, PDBD-HI, and TDBD-HI) across two informants resulted in a R2 of .984. As such, the incremental utility of a structured diagnostic interview (or any assessment method for that matter) was theoretically futile in relation to ADHD symptomatology. Furthermore, parent ratings on a symptom-based rating scale alone resulted in better fit, greater association, and greater classification than a structured diagnostic interview alone within our sample. Using either informant’s (parent or teacher) ratings of inattentive or hyperactive/impulsive symptoms resulted in rates of agreement ranging from 87.3% to 93.6% with diagnoses derived from a comprehensive, “gold standard” assessment of ADHD. Importantly, no other combination of measures, methods, or raters was able to correctly classify more participants, result in greater fit, or account for greater variance than the combined parent and teacher symptom-based ratings. Surprisingly, this outcome remained true even when combining the ratings on a parent-completed structured diagnostic interview and teacher-completed rating scales. This finding suggests that the information provided by symptom-based rating scales not only uniquely informs diagnosis, but is also essential in the assessment of ADHD and likely renders the symptom information provided by a structured diagnostic interview redundant.
It is noted that the maximal rate of agreement with consensus diagnoses of ADHD-C and ADHD-HI was driven exclusively by parent ratings of inattentive and hyperactive/impulsive symptoms and teacher ratings of hyperactive/impulsive symptoms of ADHD. We argue that this finding is partially attributable to our sample as only children with ADHD-combined type or ADHD-hyperactive/impulsive type were included within the study. We note that teacher ratings of inattention were independently associated with ADHD diagnosis and contributed to significantly greater incremental fit in regression models excluding any one of the three scales noted above.
Our findings also provide continued support for the use of multiple informants in the assessment of ADHD. Following DSM-IV-TR guidelines and near universal agreement affirming the collection of information in multiple settings, the importance of teacher ratings in diagnostic classification was supported in this study. In fact, despite the notable robustness of models including parent ratings alone, the addition of teacher ratings contributed significantly to greater model fit. We also argue for the continued inclusion of teacher ratings of inattention given their independent association with an ADHD diagnosis in this study as well as the litany of research documenting their predictive utility (Power, Doherty, et al., 1998; Power et al., 2001). Furthermore, as our sample included only children with ADHD-C or ADHD-HI, the inclusion of teacher ratings of inattention may be even more critical in assessing for ADHD-I.
Finally, this study demonstrated the clinical utility of behavioral rating scales in the assessment of ADHD. Our examination of clinically relevant diagnostic algorithms provided empirical support for their use as a proxy for the logistic regression models. In fact, our data indicated that a diagnostic algorithm using behavioral rating scales alone was able to correctly classify all 267 of our participants. Although this finding is likely inflated due to overlap between predictors and criterion, it provides strong support for previous research suggesting that little to no additional information is added to the prediction of an ADHD diagnosis beyond that provided by behavioral rating scales (Pelham et al., 2005; Power et al., 2001; Simonsen & Bullis, 2007). Although it is theoretically feasible for clinicians to use statistical models to inform diagnostic decision making, in reality this process is likely untenable. We contend that this examination of diagnostic algorithms moves incremental validity research into the realm of clinical relevance through considering cost-efficiency, practicality, and real-world utility (Johnston & Murray, 2003). Specifically, as suggested by previous research (Power et al., 2001; Sayal, Letch, & Abd, 2008; Simonsen & Bullis, 2007), our findings provide support for the use of a multiple-gating procedure in the assessment of ADHD. Using a multiple-gate procedure would likely decrease the expenditure of unnecessary resources required for “gold standard” assessments (especially those requiring a structured diagnostic interview). For example, children likely to meet criteria for ADHD may require only parent and teacher rating scale data and a semistructured clinical interview (assessing developmental, social, academic, and family functioning as well as age of onset and impairment). This approach would allow for other resources to be applied toward treatment planning and intervention. Similarly, children unlikely to meet criteria for a diagnosis of ADHD may be identified using efficient behavioral rating scales without requiring further parent, teacher, child, or clinician resources.
Strengths and Limitations
This study had several notable strengths and limitations. First, determinations about group membership (e.g., diagnostic status) were based on a consensus decision agreed on by pairs of psychologists after independently reviewing information provided by a comprehensive evaluation strategy that included parent and teacher rating scales, a structured diagnostic interview, a full cognitive and achievement battery, child self-report measures, and a clinical interview assessing developmental, social, and academic functioning. Despite the presumed validity of this “gold standard” diagnostic procedure, the results may have differed if an alternative diagnostic procedure was used as the criterion for an accurate diagnosis. Furthermore, this is the first study to examine the relative incremental utility of each method within a comprehensive diagnostic procedure providing an accurate measure of their actual use by multiple diagnosticians in informing the diagnostic decision. However, further research is warranted examining the incremental utility of assessment methodology, which is independent of the criterion diagnosis.
We raise three issues for discussion and possible consideration in future research regarding the DISC interview. First, the ADHD module on the DISC interview was presented at the end of the structured interview, which was preceded by multiple behavioral rating scales and a semistructured clinical interview reviewing primary concerns including onset, frequency, and intensity of problem behaviors. Given past findings regarding symptom attenuation within structured interviews (Jensen & Edelbrock, 1999; Piacentini et al., 1999), it is possible that the frequency of symptom endorsement declined over the course of the structured interview in this study. However, the DISC interview is structured so that the ADHD module is administered toward the end of the interview increasing risk for attenuation. Future studies should consider administration of the ADHD module first to eliminate this concern. Second, given the criterion necessary for a symptom to be noted as present on the DISC-IV (present across settings for at least 6 months), it is possible that some symptoms rated as “pretty much” a problem on the DBD Rating Scales may not have been endorsed on the DISC-IV. The assumption of equivalence between ratings on symptom-based rating scales and their relative endorsement on a structured diagnostic interview should be examined further. Finally, this study did not examine the utility of the DISC-IV interview in identifying and/or diagnosing other disorders often comorbid with a diagnosis of ADHD. The DISC-IV likely would exhibit greater clinical utility in assessing for comorbid conditions in children with ADHD.
Although this study included a significant percentage of females (22.5%), there were neither enough girls nor enough minority children to permit comparative analysis for subsets of the sample. Future research should examine whether the incremental utility and the usefulness of a diagnostic algorithm in the assessment of ADHD is consistent across sex and racial/ethnic backgrounds.
Implications for Research, Policy, and Practice
In summary, the goals of this study were to examine the actual, unique contributions of universally recommended assessment methods in a comprehensive, “gold standard” assessment of ADHD. In relation to these goals, this study demonstrated the independent contributions of behavioral rating scales and a structured diagnostic interview in the assessment of ADHD. In addition, we demonstrated the relative incremental utility of these methods across informants in the prediction of a diagnosis of ADHD. As such, this research challenged the practice of using a structured diagnostic interview, in addition to behavioral rating scales, as an efficient and incrementally valid method of assessment. Of note, a structured diagnostic interview failed to contribute unique diagnostic information beyond that collected by much more efficient rating scales. In addition, a less time-intensive semistructured clinical interview regarding developmental, social, academic, and family functioning would provide remaining diagnostic information (e.g., age of onset). As an understanding of the incremental utility of assessment methodology is critical in bridging the gap between laboratory and clinic-based settings, we argue that future research should incorporate similar approaches to identify the most ecologically valid, cost-effective methods in assessment. Furthermore, this study provided a clinically relevant, efficient method of integrating information from behavioral rating scales completed by multiple informants in the assessment of ADHD. Although these findings may not be as robust in settings where complex comorbidities are more common, future researchers are strongly encouraged to examine strategies that facilitate the development of valid, efficient diagnostic algorithms that identify ADHD efficiently in a manner that is easily generalizable to a clinic-based setting.
Although the results of this study support the practice of requiring multiple informants’ reports of ADHD symptoms, the robust predictive utility of parent ratings alone suggests their use within a multiple-gating procedure to maximize efficiency and inform the need for a more intensive assessment (Pelham et al., 2005; Power et al., 2001; Simonsen & Bullis, 2001). However, as this study was not designed to identify optimal cut points, additional research verifying the use of such strategies and their respective clinical utility (e.g., incremental and ecological validity) in ruling in or out a diagnosis of ADHD is needed. Regardless, this study provides robust empirical support for the relative inefficiency and provision of redundant information of a structured diagnostic interview while providing a framework for the examination of incrementally valid, cost-effective assessment procedures.
Footnotes
Acknowledgements
The authors acknowledge the invaluable assistance of Daniel A. Waschbusch, Stephanie H. McConaughy, Nina M. Kaiser, Elizabeth A. Hurt, Meghan Tomb, Julia D. McQuade, Kate Linnea, and the families participating in the research study.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: National Institute of Mental Health grant number R01MH065899. The views expressed in this paper are solely those of the authors, and do not necessarily reflect the views of the National Institute of Mental Health.
