Abstract
Sibley, Coxe, and Molina provide a thoughtful discussion of the implications of our study and highlight important future directions in this line of work. They helpfully amplify several themes that space did not allow discussion of in our article. In particular, they correctly emphasize the importance of theoretical as well as statistical considerations in model selection. We also agree that clinical tests of sensitivity and specificity, taking into account different base rates and types of samples, are essential before a final algorithm would be ready for dissemination. However, we are not convinced that such tests should be limited to populations of individuals with attention-deficit/hyperactivity disorder (ADHD). Rather, they should include those with and without diagnosed ADHD in order to provide comprehensive tests of reporter sensitivity and specificity across the entire continuum of ADHD symptomatology and in relation to different populations, including other disorders and typically developing populations.
Refinement of attention-deficit/hyperactivity disorder (ADHD) assessment in young adults is a critically important direction for future work. We appreciate Sibley, Coxe, and Molina’s thoughtful response to our article, which highlights important additional considerations and future directions in this line of work. We are in full agreement with most of the points raised in their response. In particular, we agree that the complex structure and inter- (and intra-!) individual heterogeneity of individuals of all ages with the disorder complicate examination of structure, models, and diagnosis of the disorder. We debated this issue extensively when selecting a model for examination. Based on the theoretical construct in DSM-5 of ADHD as coherent syndrome characterized by two, related symptom domains of inattention and hyperactivity-impulsivity, we chose not to focus here on individual items. In the dominant model of ADHD, individual items are implicitly viewed as indicators of latent inattentive and hyperactive-impulsive symptom domains or, alternatively, of the entire ADHD syndrome. An individual item approach is, of course, another way to go, and we are also interested in that approach. But, as Sibley et al. point out, analytic decisions must be made. Although we stand by our model choice based on theoretical considerations and study goals, we certainly encourage others to examine different models based on alternative theoretical and study goal considerations so the field can collectively evaluate the most effective models for different purposes.
Sibley et al. compare our model to an alternative model proposed by Bauer et al. (2013). They questioned why our model did not allow for item-specific agreement between raters, whereas Bauer et al. examined such item-level rater agreement and found evidence for such agreement in ratings of negative affect. We were well aware of this option but deliberately excluded examination of item-specific agreement between raters in our final model based on theoretical considerations (discussed above) and also due to data reported in a prior version of this paper seen by the reviewers, as well as in a related, published article in a childhood (vs. adult) sample (Martel, Schimmack, Nikolas, & Nigg, 2015). Those data showed—unlike negative affect—little evidence for item-specific rater convergence for ADHD. Specifically, in an earlier version of this article seen by the reviewers, we empirically tested rater agreement for individual ADHD items and concluded,
Only some of the unique, or residual, variances across the same self- and peer-rated symptom items were significantly correlated, suggesting that most of the agreement among raters was occurring at the latent level and that there is little agreement in the severity of specific symptoms.
We published similar findings in a childhood sample using a modified model that allowed item-level agreement between raters to be modeled (Martel et al., 2015).
Bauer et al. (2013) likely found strong agreement between specific items in their model of negative affect because their model assumed that agreement between raters is either due to a single common factor or item-specific agreement; that is, they did not allow subfactors. In contrast, our model allowed for agreement in two distinct ADHD factors (i.e., inattention and hyperactivity-impulsivity) with well-established discriminant validity. Had we, like Bauer et al. (2013), specified only one factor superordinate, we likely would have found more item-specific agreement, but this apparent item-level agreement would be explained by actual agreement on severity of the two latent subfactors of inattention and hyperactivity-impulsivity, not actually the items.
Our results do not imply, and we do not claim, that ADHD can be fully understood with a two-factor model or that there is not syndromal heterogeneity within ADHD at the symptom item level. However, our data suggest that, after accounting for the general and specific symptom domain factors, individual items cannot be interpreted as symptoms providing unique information about ADHD, at least for symptom checklists like the one utilized in the current study. This is why we warn against use of scoring algorithms that equate individual items with symptoms and suggest that new diagnostic instruments need to be developed and validated to provide valid and unique information at the level of individual ADHD symptoms.
A further point of note is that we chose a model in which the two specific symptom domain latent factors were orthogonal to the general ADHD factor, based on analytic and theoretical considerations. The alternative, to allow correlation between the general and specific factors, would make the model difficult to interpret, even if we could overcome problems with fit and possible convergence problems. An orthogonal model has the advantage of assuming that the specific factors of inattention and hyperactivity-impulsivity capture additional variance over and above a general ADHD syndrome factor. A correlated model makes a different assumption and fits a different conception of ADHD, in line with a second-order conceptualization of ADHD symptom domains as covarying with a general ADHD factor. Here, we considered the orthogonal bifactor model, as recommended (e.g., Chen, West, & Sousa, 2006), to be more straightforward and theoretically meaningful.
We fully agree with Sibley et al. about the importance of testing the sensitivity and specificity of clinical algorithms before any attempt at deployment. We did not intend here to propose a ready-for-the-field algorithm, but rather to outline an approach to creating such algorithms. This is in contrast to a similar article we did on integration of multiple informants in diagnosis of childhood ADHD in which we advocated a specific averaging approach to multiple informant integration (based on a different pattern of results; see Martel et al., 2015; Martel, Markon, & Smith, 2017). There, we provided sensitivity and specificity tests of this approach. Here, we believed such tests to be premature. We also agree that other statistical or even mathematical approaches may be useful for evaluating multiple informant rating approaches, including person-level approaches, receiver operator characteristics curves, and network analysis (Martel et al., 2017).
Our only real disagreement with Sibley et al. lies with the nature of sampling that should be done. They suggest that samples limited to those with ADHD are most useful for testing approaches to integration of multiple informant ratings in ADHD. We are not convinced, due in part to the importance of being able to evaluate sensitivity and specificity in the same settings as rater agreement. As they point out, over- AND under-reporting of ADHD symptoms are both problems in adult ADHD populations. Furthermore, due in part (though not entirely) to the five-symptom diagnostic threshold, people move in and out of the diagnostic category over time (Lahey, Pelham, Loney, Lee, & Willcutt, 2005). Finally, prior work using multiple samples and data analytic strategies suggest that ADHD is best modeled as a continuum versus a diagnostic category (Haslam et al., 2006; Larsson, Anckarsater, Råstam, Chang, & Lichtenstein, 2012; Marcus & Barry, 2011; Stergiakouli et al., 2015). While more confirmation of this structure remains needed, at present, recruiting those both with and without formal ADHD criteria would seem necessary. Furthermore, alternative samples are obviously needed for other purposes (e.g., discriminating ADHD from other clinical conditions). Our sample almost equally included those with and without ADHD, including subthreshold cases. Therefore, it provides potential insight into the structure of ADHD symptomatology as the continuum that it seems to be. Yet as we noted in our study limitations and as pointed out by Sibley et al., we—of course—agree that our results, like practically all studies, need not only replication but also generalizability evaluation in other samples (e.g., epidemiological; forensic).
The convenience sample used here is not without merit, however; it includes very well characterized ADHD (not usually available in epidemiological samples), avoids clinic referral biases (as in forensic samples), and enables us to illustrate the promise of this approach, which now requires replication and extension. At the same time, it has limitations, including possible volunteer bias and possible exclusion of the severe end of spectrum. The important concept here is to look at results across multiple sample types, but probably not limited to samples of ADHD.
In sum, we commend Sibley et al. for their thoughtful contribution to the discussion and concur with them that multiple studies in multiple samples using multiple data analytic strategies are necessary for next-generation clinical algorithm development (Martel et al., 2017). We hope our study continues to spur such work.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
