Abstract
The present study aimed to investigate the measurement invariance across age, gender, clinical status, and informant of the Attention-Deficit/Hyperactivity Disorder Rating Scale–IV (ADHD-RS-IV) Home and School versions. The participants were 1,106 Romanian children and adolescents (mean age = 12.74 years, standard deviation = 2.84, age range 6-18 years). Both parents and teachers assessed ADHD symptoms. The factorial structure of the scale was assessed using confirmatory factor analysis, and measurement invariance was assessed using multigroup confirmatory factor analysis. The results supported the reliability of the ADHD-RS-IV, with high internal consistency coefficients for both versions. Confirmatory factor analysis validated a two-factor model. Multigroup confirmatory factor analysis confirmed the measurement invariance of ADHD-RS-IV across age, gender, clinical status, and informant. ADHD-RS-IV had good psychometric properties in a sample of Romanian children and adolescents. It is a reliable instrument given its strong invariance. Implications for evidence-based assessment of ADHD are discussed.
Keywords
Attention deficit/hyperactivity disorder (ADHD) is a condition mostly characterized by inattention, hyperactivity, or impulsivity (American Psychiatric Association [APA], 2013). ADHD is one of the conditions most frequently encountered in children and adolescents, with approximately 3% to 5% of youths suffering from ADHD (Polanczyk, de Lima, Horta, Biederman, & Rohde, 2007; Polanczyk, Salum, Sugaya, Caye, & Rohde, 2015). These children and adolescents are at risk for school failure, learning disorders, behavior problems, and problematic social relationships (Barkley, 1997; Hinshaw, 1994; Sexton, Gelhorn, Bell, & Classi, 2012). Moreover, the burden of disease is large for families of youths diagnosed with this condition (Gupte-Singh, Singh, & Lawson, 2017).
When assessing children and adolescents suspected of having ADHD, it is essential for clinicians to use reliable and validated instruments that take into consideration the prevalence, severity, and high number of socio-emotional and cognitive difficulties associated with this disorder. Evidence-based assessment in ADHD includes structured diagnostic interviews with the child, parents, and teachers; completion of ADHD rating scales by the parents and teachers; and direct observation of behavior at school and in clinical testing situations (Barkley, 2014; DuPaul & Stoner, 2014; Pelham, Fabiano, & Massetti, 2005). Parent and teacher ratings of child behavior are important; in fact, Pelham and collaborators (2005) show that when such reports are obtained from parents and teachers, there is no increase in the incremental validity given by the use of structured interviews. Therefore, ADHD rating scales filled in by parents and teachers represent easy to administer instruments through which ADHD can be detected, at minimal cost, without using extensive interviews, which require more resources from clinicians and parents and are also expensive and not practical for those cases where multiple assessments are conducted at different time periods.
The importance of the information provided by family members or teachers in establishing an ADHD clinical diagnosis has been demonstrated in many studies (Glascoe, 2000; Mulhern, Dworkin, & Bernstein, 1994; Young, Davis, Schoen, & Parker, 1998). Parents hold information regarding not only the relevant antecedents in the disorder’s history but also the evolution of the behavior over time, which together can offer an integrated image of the child’s psychopathology (Mash & Terdal, 1997). Frequently, this information allows parents to identify other relevant attributes that may remain unidentified during clinical testing (Dewey, Crawford, Creighton, & Sauve, 2000). Teachers can also provide relevant information concerning the child’s behavior at school, including the specific manner of interaction with colleagues and peers (Mash & Terdal, 1997; Miller, Koplewicz, & Klein, 1997). Related studies show that a teacher’s assessment has higher reliability than the one offered by parents while being at the same time more sensitive to hyperactive behavior (Barkley, 2014).
Because of its simple structure and application method, the ADHD Rating Scale–IV (DuPaul, Power, Anastopoulos, & Reid, 1998) is an instrument proven to be sensitive in intervention studies where repeated application is used to monitor behavioral changes (DuPaul & Stoner, 2014). ADHD-RS-IV allows clinicians the opportunity to obtain data from both parents (ADHD-RS-IV Home version; DuPaul, Anastopoulos, et al., 1998) and teachers (ADHD-RS-IV School version; DuPaul et al., 1997) regarding the frequency of each characteristic ADHD symptom according to established Diagnostic and Statistical Manual of Mental Disorders, fourth edition (DSM-IV; APA, 2000) criteria. Regarding the factorial structure of ADHD-RS-IV results, exploratory factor analysis tested a one-factor, two-factor, three-factor, and, more recently, modified two-factor solution (Döpfner et al., 2006; DuPaul, Anastopoulos, et al., 1998; DuPaul et al., 1997; Martel, Von Eye, & Nigg, 2010; Sturm, McCracken, & Cai, 2017). Results of confirmatory factor analysis (CFA) sustained a two-factor structure of the scale, with inattention and hyperactivity being the two factors considered in a Chinese sample (Su et al., 2015), Japanese samples (Takayanagi et al., 2016; Tani, Okada, Ohnishi, Nakajima, & Tsujii, 2010), an Icelandic sample (Magnússon, Smári, Grétarsdóttir, & Prándardóttir, 1999), participants from 10 European countries (Döpfner et al., 2006), and another multinational study comprising participants from several European countries and Australia, Israel, and South Africa (Zhang, Faries, Vowles, & Michelson, 2005).
Although behavioral assessment scales developed in one country often are applied in other countries, it is important to understand how the adapted version of the scale functions in a specific country (Ivanova et al., 2007). At first, it was believed that the expression and development of psychological disorders were largely universal attributes unaffected by ethnicity (Marsella & Kameoka, 1989). Later studies, however, have shown that cultural bias is an issue with ADHD testing (Reid, 1995). In this context, and according to international norms involving the translation/adaptation of psychological testing instruments (Hambleton & Patsula, 1998), the main concern in evidenced-based assessment is how to validate the adapted scale to the culture in which it was initially created. According to a recent systematic review conducted on the existent cross-cultural invariance of assessment instruments developed for children and adolescents, where the ADHD-RS-IV School version was also included, there is a lack of strong evidence regarding the appropriateness of this scale for use in cross-cultural studies (Stevanovic et al., 2017).
Measurement invariance, defined as the psychometric property of a test that shows equivalence in the latent variable analyzed (Meredith, 1993; Vandenberg & Lance, 2000), is critical to establish before using a scale, either in research or in clinical practice, given the fact that errors can appear in item understandings. Therefore, a precondition of valid comparisons between different groups (e.g., boys and girls, children and adolescents, clinical and nonclinical samples) is establishing that the instrument used is invariant—that is, its items have the same meaning across different groups. Previous studies showed that several items from other ADHD rating scales can function differently across boys and girls, younger and older children (Makransky & Bilenberg, 2014), and children with or without a diagnosis of ADHD (Li, Reise, Chronis-Tuscano, Mikami, & Lee, 2016). However, these results are mixed, with more recent findings showing no differential item functioning by age and gender (Sturm et al., 2017).
The Present Study
Most of the research so far on the psychometric properties of the ADHD-RS-IV has been conducted with American samples of children and adolescents. There are just a few studies that used the ADHD-RS-IV in European countries; however, in these studies, the instrument was either physician (Döpfner et al., 2006) or clinician (Zhang et al., 2005) rated based on the semistructured interviews conducted with parents of children with ADHD. Therefore, there is limited research from east European countries on the psychometric properties and measurement invariance of the ADHD-RS-IV Home and School versions. Romania is a collectivistic country; therefore, differences between Romanian parents’ and teachers’ interpretations of the ADHD-RS-IV items could emerge as compared with parents and teachers from individualistic countries. These differences could be explained by parenting behavior and expectations/attitudes regarding children’s behavior, which were profoundly shaped by the communist culture (before 1989). According to Davidov, Dülmer, Schlüter, Schmidt, and Meuleman (2012) measurement invariance, in particular scalar noninvariance, represents one of the most serious threats to cross-cultural research. The literature on adult samples shows that there are significant differences in countries’ norms on gender equality between western Europe and central and eastern Europe (Weziak-Bialowolska, 2015). Moreover, there are differences between central and eastern European countries as well, with Romania being one of the lowest-scoring countries regarding attitudes toward gender equality.
Given the importance of parent and teacher ADHD rating scales over interviews (Pelham et al., 2005), through adaptation and validation of the instrument into the Romanian language, we ensured that Romanian children with ADHD can be involved in international clinical trials and results would not be affected by errors in item understanding. Furthermore, given the limited research conducted on measurement invariance of the instruments that assess psychopathology in children, it is highly important to add to the existing literature data on the measurement invariance of one of the most used scales assessing ADHD. As far as we know, no study has investigated so far the measurement invariance of the ADHD-RS-IV across age, gender, clinical sample, and informant, although such research exists on other instruments used in the assessment of ADHD (see, e.g., evidence for a modified ADHD-RS in Danish participants; Makransky & Bilenberg, 2014) that show differential item functioning.
Furthermore, given the variability in the ADHD prevalence rate across countries, which according to systematic reviews and other researches (Polanczyk et al., 2007; Thomas, Sanders, Doust, Beller, & Glasziou, 2015; Valo & Tannock, 2010; Willcutt, 2012) resulted rather from methodological differences and not from differences due to geographical locations, we aim to standardize (providing norms for scores interpretation) an instrument easy to administer and score that can be used in prevalence studies. So far, the prevalence of ADHD in Romanian children and adolescents has not been established.
Taking into consideration all the aspects mentioned above, the first aim of the present study was to analyze the reliability (internal consistency and interrater correlation) of the ADHD-RS-IV, both Home and School versions. Second, we aimed to investigate the factorial structure of the Romanian version of the ADHD-RS-IV Home and School versions. Third, we aimed to answer the next research question, regarding the scale’s measurement invariance: Does the ADHD-RS-IV function similarly across gender, age, and clinical status of the child for both parent and teacher ratings? Furthermore, we wanted to investigate the differences between girls and boys, children and adolescents, clinical and nonclinical samples in the ADHD latent variable. Finally, our fifth objective was to present normative data for the Romanian version of the ADHD-RS-IV Home and School versions.
Methods
Participants
Overall, the study sample included the parents and teachers of 1,106 Romanian children and adolescents. The youths’ age range was between 6 and 18 years; the mean (M) age was 12.74 years, and the standard deviation (SD) was 2.84. Gender distribution was relatively balanced, with 47.4% boys and 52.6% girls participating. The home environment distribution also was comparatively balanced, with 57.3% of participants living in urban areas and 42.7% in rural areas. The sample comprised nonclinical and clinical children. Nonclinical participants (N = 1,046) were selected from Romanian schools, based on a stratified random sampling procedure. Based on this sampling method, we obtained a representative national school-based sample of Romanian children in terms of gender, age, and school grades. The educational staff completing the questionnaires were mostly teachers (96.8%), with the rest being tutors (0.7%), educational psychologists (0.9%), advisors (0.9%), or those in other roles (0.7%). The time spent by the educational staff with the students was between 1 and 40 hours a week (M = 10.27 hours, SD = 8.09). This difference in the degree of acquaintance with the children resulted from the structure of the Romanian system of education, where teachers spend more time with children in primary school (teachers spend normally 4 hours per day with children in primary school and even more hours a day with children in the afterschool system) whereas in secondary school the amount of time spent with children is less, depending on the subject taught (schoolmasters spend a minimum of 1 hour per week depending on the subject they teach: e.g., maths, 4 hours per week; music, 1 hour per week).
Most of the staff members reported either a moderate degree of acquaintance (we included in this category teachers spending between 4 and 6 hours per week with the student) (53.1%) or a very good one (teachers spending between 7 and 10 hours per week with the student) (40.8%). Those at the extreme ends represented a small number who reported either a very high degree of acquaintance (0.5%) with the child or a very low one (5.6%). Clinical participants (N = 60), with a primary diagnosis of ADHD, were selected from an infant psychiatric unit in Romania where they were referred for diagnosis and treatment for the first time. Children included in the clinical group were diagnosed by a child psychiatrist according to the criteria from the International Statistical Classification of Diseases and Related Health Problems, 10th Revision (World Health Organization, 1992). No children from the clinical group were under treatment (neither psychological nor pharmacological). We found a significant difference between the clinical and nonclinical groups regarding their age, t(1104) = 13.89 (p = .001), the nonclinical group having a higher mean age, M = 13.01 (SD = 2.68) than the clinical group, M = 8.31 (SD = 1.58). As far as sex differences were concerned, we found a significant association between clinical status and sex category, χ2(1) = 30.21, p = .001, with the clinical group including more boys (81%) than the nonclinical group (45.3%).
Measures
Demographics
Parents and teachers completed a demographic questionnaire regarding the children and adolescents assessed (age and gender) and their residence (urban, rural).
ADHD Symptomatology
The instrument used in our study was the ADHD-RS-IV, proposed by DuPaul, Power, et al. (1998). The ADHD-RS-IV has two identical forms, one for parents (ADHD-RS-IV Home version) and one for teachers (ADHD-RS-IV School version), that assess children’s behavior at home and in school, respectively. Each form contains 18 items corresponding to the ADHD symptoms described in DSM-IV (APA, 2000). Behavior assessment by parents and teachers was done on a 4-unit scale, with 0 = Never or Rarely, 1 = Sometimes, 2 = Often, and 3 = Very often. The score structure is related to the way DSM-IV describes this disorder. Three scores were calculated based on the results: (1) the hyperactivity and impulsivity (HI) dimension score, (2) the inattention (IA) dimension score, and (3) the total score (TS). The TS was calculated by adding the answers for each item; in this study, the TS ranged between 0 and 54. To address possible response bias, IA symptoms were designated as odd-numbered items and HI symptoms as even-numbered items on the 18-item scale.
Procedure
The Romanian ADHD-RS-IV was adapted to norms associated with psychological testing instruments (Geisinger, 1994; Hambleton, 1994; Hambleton & Patsula, 1998). Dyads of translators did the translation and backward translation of the two scales. The backward translation aimed at the conceptual equivalent of each item, and English items were translated in the most relevant way. Each translator was a health professional with at least 7 years of translating experience, familiar with the terminology of the area covered by the instrument (i.e., child psychopathology) and knowledgeable about the English language (based on a certificate of competence after graduation from a bilingual high school), even if the mother tongue was Romanian. In this study, each parent signed an informed consent form to allow his or her children’s data to be collected and analyzed; children older than 7 years also provided written assent. The children’s assessment was performed by one parent (on the ADHD-RS-IV Home version) and by teachers (on the ADHD-RS-IV School version).
Data Analysis
Univariate and bivariate descriptive indicators, between-groups statistical comparisons, percentile rank calculation for each normed group, and ordinal Cronbach’s alpha (Dunn, Baguley, & Brunsden, 2014) were conducted using the IBM SPSS for Windows (Version 23) statistical software extended by the “userfriendlyscience” R 3.1.3 package. CFA and multigroup confirmatory factor analysis (MGCFA) were conducted using Mplus Version 7.4 (L. K. Muthén & Muthén, 1998–2015). The construct validity of the ADHD-RS-IV was investigated in a series of exploratory factor analytical studies (DuPaul, Anastopoulos, et al., 1998; DuPaul et al., 1997). In this study, we conducted a CFA for the ADHD-RS-IV (both Home and School versions) on a nonclinical representative national Romanian sample. Using CFA, we tested two measurement models, a one-factor and a two-factor model, to find the best-fitting one. The two-factor model was proposed by DuPaul, Anastopoulos, et al. (1998) and DuPaul et al. (1997) and contains two dimensions: IA and HI. Both forms of the questionnaire, Home and School versions, were analyzed to assess their factorial structure.
Similarities in the relationships between ADHD-RS-IV latent constructs and scale items across gender (male vs. female), age- (≤12 vs. >12 years), and status (nonclinical vs. clinical) groups were tested through MGCFA using Theta parametrization. Given the ordered-categorical nature of the scale, we considered that factor loadings and thresholds jointly define item functioning; as a consequence, they were constrained and freed together, reducing the number of measurement models tested through MGCFA (three-step approach) (Bowen & Masa, 2015).
The application of normal theory–based estimation procedures to ordered-categorical data (such as those provided by the ADHD-RS-IV) frequently results in biased parameter estimates and inaccurate statistical significance tests (because of biased standard errors) (B. Muthén & Hofacker, 1988). The ADHD-RS-IV uses a 4-point Likert-type scale; as a consequence, CFA and MGCFA were carried out using robust weighted least squares (WLSMV—weighted least squares with means and variance adjusted). WLSMV was designed specifically for ordered-categorical data; its parameter estimation is based on a polychoric correlation matrix (Beauducel & Herzberg, 2006; Rhemtulla, Brosseau-Liard, & Savalei, 2012).
Evaluation of the model fit of each tested measurement model (CFA or MGCFA) to the data was based on several fit indicators: Satorra-Bentler chi-square (SBχ2; Satorra & Bentler, 2001), comparative fit index (CFI; Bentler, 1990), Tucker–Lewis index (TLI; Bentler & Bonett, 1980), and root mean square error of approximation (RMSEA; Steiger, 1990). Local model misfit (e.g., low factor loadings) and interpretability of parameter estimates (e.g., eventually Heywood cases) were used as additional criteria for model fit evaluation. Criteria for a good model fit were nonsignificant SBχ2, CFI > .95, TLI > .95, and RMSEA < .05 (Hu & Bentler, 1999). Acceptable model fit was considered when CFI > .90, TLI > .90, and RMSEA < .08 (Browne & Cudeck, 1993; Marsh, Hau, & Wen, 2004).
Nested models fit comparison (within CFA or MGCFA) was performed using the Mplus DIFFTEST procedure, which uses robust standard errors and adjusted chi-square (Lubke & Muthén, 2005). A nonsignificant delta of SBχ2 (ΔSBχ2) was used as an indicator of measurement invariance. Given that large sample size has a major impact on chi-square (Hoelter, 1983), the decision on model fit comparison was made on other difference fit indicators such as ΔCFI, ΔRMSEA, and ΔTLI. Regarding ΔCFI and ΔRMSEA, a change of ≥.010 in CFI, supplemented by a change of ≥ .015 in RMSEA would indicate noninvariance (a significant worsening of model fit) (Cheung & Rensvold, 2002; Dimitrov, 2010). According to Little (1997) ΔTLI ≤ .05 would be interpreted as a nonsignificant change of model fit.
Results
Our results are presented in two sections. The first section presents descriptive statistical results (M and SD) and reliability estimates. ADHD-RS-IV reliability was assessed through analysis of internal consistency (using ordinal Cronbach’s alpha coefficient computed for each scale). The second section presents the results of CFA and measurement invariance (using MGCFA) for both parent- and teacher-rated versions of the scale.
Descriptive Statistics for Both Models and Reliability of ADHD-RS-IV
Table 1 presents descriptive statistics for both ADHD-RS-IV forms, including individual factors (IA and HI) as well as the TS for parents and teachers. The main diagonal shows the internal consistency index for each scale, estimated through ordinal Cronbach’s alpha (Dunn et al., 2014).
Descriptive Statistics for Both ADHD-RS-IV Forms, and the Internal Consistency Index (Ordinal Cronbach’s Alpha) for Each Factor (IA and HI) and TS.
Note. ADHD-RS-IV = Attention Deficit/Hyperactivity Disorder Rating Scale–IV; IA = inattention factor; HI = hyperactivity/impulsivity factor; TS = total score; M = mean; SD = standard deviation.
The results indicate high internal consistency of the ADHD-RS-IV items, in both Home and School versions. Cronbach’s alpha varied between .87 and .96. We found a remarkably lower Cronbach’s alpha for teachers’ rating, compared with parents’. Differences in the variances of the assessed samples might lead to differences between the alpha estimates; if a sample is more homogeneous, it will often lead to lower standard deviations on the items, and this will tend to lower alpha.
The interrater correlation between the scores obtained for each subscale, and TS varied between .38 and .43. Correlations of the same construct between parents and teachers were adequate, r = .43, p < .05 for IA and r = .38, p < .05 for HI. We expected to observe lower correlations when different dimensions were assessed by different evaluators. Table 1 confirmed those expectations: The correlation between teachers HI and parents IA had r = .34 (p < .05), whereas teachers IA and parents HI had r = .27 (p < .05).
Confirmatory Factor Analysis
Two models were examined: (1) a one-factor model, where each item loads on a single latent variable, and (2) a two-factor model, with two latent variables, where odd-numbered items load the IA factor and even-numbered items load the HI factors. Figure 1 presents the graphical representations of the recurrent model specifications of the two models. To set the latent variable scale, we used the unit loading identification method (Kline, 2005).

Conceptual diagrams of models with (a) a single latent variable and (b) two latent variables.
The fit statistics for each model, the one-factor and two-factor models (using data from the parents’ and teachers’ assessments), are presented in Table 2. Compared with the parents’ evaluations, the teachers’ data generally showed a lower fit to either measurement model. The chi-square for all the models was found to be significant, meaning that there is a significant discrepancy between the observed and reproduced variance–covariance matrices. Given the large sample size, this indicator is hard to interpret because of its oversensitivity to any discrepancy. Other fit indices, except for RMSEA, were situated in the acceptable range for the one-factor model, and the fit to the two-factor model was good. The RMSEA values for all models were in the acceptable range, showing an unacceptable high value (>0.08) only for the one-factor model for parents.
CFA Fit Indicator Values for the One-Factor Model and the Two-Factor Model.
Note. CFA = confirmatory factor analysis; df = degrees of freedom; CFI = comparative fit index; TLI = Tucker-Lewis index; RMSEA = root mean square error of approximation; CI = confidence interval; Δχ2 = chi-square difference; Δdf = degree of freedom difference.
Significant χ2 values.
The superiority of the two-factor model was also supported by the statistically significant value of Δχ2 (df = 1) = 67.88 (p < .01) and of ΔCFI = −.017 for the ADHD-RS-IV Home version. ΔTLI (−.019) and ΔRMSEA (.01) for the Home version, according to the established cutoffs, were not found to be significant. For the ADHD-RS-IV School version, the two-factor model showed a clear superiority of model fit, Δχ2 (df = 1) = 192.15 (p < .01), ΔCFI = −.046, ΔTLI = −.051, and ΔRMSEA = .045, indicating a significant improvement of the model fit after the model was respecified as a two-factor model.
Taking into account all the fit indicators and the results of the nested model comparisons, we concluded that the two-factor model has a better fit to the data than the one-factor model (see Table 2).
Table 3 presents the factor loadings for each item; again, the two-factor model proved to be more efficient than the one-factor model. All the items’ standardized loadings for the two-factor model were higher than the absolute value of .30. In general, the values of the factor loadings were higher for teachers than for parents. The explained item variance by the latent factor varied between 45% and 92% for the parents’ data and between 76% and 96% for the teachers’ data.
CFA Standardized Regression Coefficients of the ADHD-RS-IV Items and Factor Correlation.
Note. ADHD-RS-IV = Attention Deficit/Hyperactivity Disorder Rating Scale–IV; CFA = confirmatory factor analysis; IA = inattention factor; HI = hyperactivity/impulsivity factor.
Measurement Invariance
To test the measurement invariance by child age, gender, and clinical status, the two-factor model fit was tested in each subpopulation; the goodness-of-fit indicators are presented in Table 4. Following the two-step procedure (Dimitrov, 2010), only configural invariance and strong invariance were tested. As a preliminary step of testing measurement invariance, the two-factor model was tested within each invariance subgroup (child age, child gender, child clinical status, and informant) using CFA for ordinal data. The two-factor structure of the questionnaire fit the data well; although the robust SBχ2 was statistically significant, the CFI, TLI, and RMSEA values were all good or acceptable according to all the established guidelines (Hu & Bentler, 1999).
Two-Factor CFA Fit Indicators by Child Age, Gender, and Clinical Status Subgroups for the Questionnaire Form (Parents and Teachers).
Note. CFA = confirmatory factor analysis; N = sample size; CFI = comparative fit index; TLI = Tucker-Lewis index; RMSEA = root mean square error of approximation; CI = confidence interval.
Configural invariance results across age, gender, clinical status, and informant showed that both Home and School versions of the Romanian ADHD-RS-IV were best described by the two-factor structure, across all subgroups. Regarding child gender, the configural model showed good fit indicators, CFI = .96, TLI = .954, and RMSEA = .054 for parents, and CFI = .977, TLI = .974, and RMSEA = .064 for teachers, all factor loadings being significant (p < .05). Similar results were found for the configural models involving age (CFI = .963, TLI = .958, and RMSEA = .052 for parents; CFI = .978, TLI = .975, and RMSEA = .066 for teachers), clinical status (CFI = .962, TLI = .957, and RMSEA = .049 for parents; CFI = .979, TLI = .976, and RMSEA = .062 for teachers), and informant (CFI = .978, TLI = .975, and RMSEA = .063).
To establish scalar invariance, the factor loadings and thresholds were simultaneously constrained across groups. With regard to gender groups, the results show a good model fit (CFI = .966, TLI = .967, and RMSEA = .046 for parents, and CFI = .979, TLI = .98, and RMSEA = .056 for teachers). When compared with the configural model, the chi-square difference suggests that scalar invariance holds only for gender groups in the parents’ version of ADHD-IV (Δχ2 = 59.97, p > .05). But the likelihood ratio test, on which the chi-square is based, is sensitive to sample size; so we used the change in CFI to decide on the scalar invariance of the scales (Cheung & Rensvold, 2002). When subtracting the CFI of the configural model from the CFI of the scalar invariance model, we found differences ranging between 0 and .006, providing strong evidence of scalar invariance of the scale. Similar results were found for RMSEA differences; they ranged in absolute values, between 0 and .008. ΔTLI values ranged between .006 and .013, indicating the presence of scalar invariance.
The same pattern of results was found for the age, clinical status, and informant variables. The results indicated that the thresholds and factor loadings were invariant across gender, age, clinical status, and informant.
Examining measurement invariance using clinical status as a grouping variable resulted in largely unbalanced sample sizes. Unbalanced sample sizes across the groups being examined have been reported as a major problem in testing and interpreting measurement invariance (Chen, 2007). The larger group has more weight in determining estimated parameter values; as a consequence, eventual violations of invariance in the smaller group pass undetected (Kaplan & George, 1995). To reduce the probability of obtaining measurement invariance across clinical status because of the unbalanced nature of the design, we used a subsampling approach to adjust for unbalanced sample size (Yoon & Lai, 2018). According to the procedure described by the authors, we sampled from the nonclinical participants group 100 random subsamples, each of them having the same size as in the clinical group. Configural and scalar invariance analyses were run for all the resultant subsamples (Table 5). RMSEA, CFI, and TLI fit indicators were averaged across the samples, and then ΔRMSEA, ΔCFI, and ΔTLI were computed. The estimated differences in fit indices for the parents’ report were ΔRMSEA = −.002, ΔCFI = −.003, and ΔTLI = .003, and the same differences for the teachers’ report were ΔRMSEA = −.004, ΔCFI = −.002, and ΔTLI = .004. The delta values obtained sustain the measurement invariance of ADHD-RS-IV across clinical status groups.
Results of Measurement Invariance Using Multigroup Confirmatory Factor Analysis for the Parents’ and Teachers’ Samples.
Note. df = degree of freedom; CFI = comparative fit index; TLI = Tucker-Lewis index; RMSEA = root mean square error of approximation; Δχ2 = chi-square difference; ΔCFI = CFI difference; ΔRMSEA = RMSEA difference.
p < .05.
Gender, Age, and Clinical Status Differences
Given the measurement invariance of the ADHD-IV across gender, age, and clinical status we moved to test statistically the mean difference of IA and HI scores across gender, age, and clinical status for both versions of the questionnaire. For the ADHD-RS-IV Home version, we found that boys scored higher on the IA scale (M = 6.60, SD = 5.71) than girls (M = 4.56, SD = 4.19; t = 6.66, df = 1,056, p = .001, d = 0.40). The same pattern of results was found for the HI scale mean scores (M = 5.56, SD = 5.64 and M = 3.83, SD = 3.86; t = 5.93, df = 1,056, p = .001, d = 0.36). Regarding age differences, we found that younger children scored higher on both IA and HI scales. The mean IA for children under 12 years was M = 6.49 (SD = 5.83), whereas the same score for older children was M = 4.73 (SD = 4.17) (t = 5.70, df = 1,055, p = .001, d = 0.34). For the HI scores, we found significant differences between age-groups: M = 5.61 (SD = 5.81) for under 12 years and M = 3.84 (SD = 3.69) for over 12 years (t = 5.98, df = 1,055, p = .001, d = 0.36). The clinical sample scored higher on the IA scale: M = 15.74 (SD = 5.82) for the clinical sample and M = 4.89 (SD = 4.27) for the nonclinical sample (t = 18.91, df = 1,056, p = .001, d = 2.12); the same pattern was found for the HI scale: M = 15.52 (SD = 6.28) for the clinical sample and M = 3.97 (SD = 3.85) for the nonclinical sample (t = 21.84, df = 1,056, p = .001, d = 2.29).
Generally, teachers’ evaluation followed the same trend for each group’s differences. Male IA mean scores were found to be higher than female scores: M = 7.75 (SD = 6.38) for males and M = 4.4 (SD = 4.65) for females (t = 9.947, df = 1,085, p = .001, d = 0.59). Regarding HI scores, we found that M = 5.89 (SD = 6.05) for males and M = 2.64 (SD = 3.67) for females (t = 10.26, df = 1,085, p = .001, d = 0.61). Once again, children younger than 12 years registered higher mean scores on both IA, M = 6.99 (SD = 6.57) versus M = 5.2 (SD = 4.93) for children older than 12 years (t = 5.10, df = 1,085, p = .001, d = 0.30), and HI, M = 5.31 (SD = 6.15) versus M = 3.29 (SD = 4.62) for children older than 12 years (t = 6.20, df = 1,085, p = .001, d = 0.37), scales. Finally, we found that the clinical sample scored higher on both IA, M = 15.29 (SD = 6.19) versus M = 5.43 (SD = 5.26) for the nonclinical sample (t = 14.18, df = 1085, p = 0.001, d = 1.71), and HI, M = 13.74 (SD = 8.09) versus M = 3.60 (SD = 4.66) for the nonclinical sample (t = 15.73, df = 1085, p = .001, d = 1.53), scales.
Regarding across-informant differences, there were no differences between parents and teachers in evaluating IA (t = 1.937, df = 2,140, p = 0.053, d = 0.08), whereas we found a significant difference in evaluating HI (t = 2.094, df = 2,140, p = .036, d = 0.09). Despite the statistical significance of one evaluation compared with the other, when effect size is taken into account, both can be included in the small effect size category.
Normative Data
Tables 6 and 7 illustrate the normative data for the parent and teacher ratings of ADHD symptomatology, respectively. Normative data for child gender by child age category (≤12 vs. >12 years) were computed according to each responder (parent vs. teacher). Furthermore, cutoff scores for the 80th, 90th, 93rd, and 98th percentiles are presented to be used for assessment purposes (e.g., screening vs. identification) (Power, 1992).
Normative Data for Parent Ratings.
Note. M = mean; SD = standard deviation; N = sample size.
Normative Data for Teacher Ratings.
Note. M = mean; SD = standard deviation; N = sample size.
Discussion
ADHD is a condition prevalent in children and adolescents (Polanczyk et al., 2007, 2015); it is associated with negative consequences for youths and their families (Barkley, 1997; Hinshaw, 1994; Sexton et al., 2012), as well as with a high economic burden (Gupte-Singh et al., 2017). To identify and offer effective treatment for this condition, the clinical diagnosis should be based on reliable instruments with sound psychometric properties. The ADHD-RS-IV (DuPaul, Power, et al., 1998) is one of the most frequently used scales and has been translated into several languages and used in several cultures; however, most of the research investigating the psychometric properties of this scale has been conducted with samples from Western countries. Research on the measurement invariance of the instrument is limited, as is the case with other instruments used in the assessment of psychopathology in youths (Stevanovic et al., 2017). The aim of the present study was to investigate the psychometric properties of the Romanian version of ADHD-RS-IV, both Home and School versions, in a national representative sample of children as well as in a clinical sample of children diagnosed with ADHD. Next, we aimed to investigate its model fit and measurement invariance across child age, gender, and clinical status, and informant. Our results indicated adequate reliability of the scale (e.g., internal consistency). Cronbach’s alpha values for the parent version varied between .79 and .89. These values are similar to the ones reported in two multinational studies investigating the reliability of the Home version (Döpfner et al., 2006; Szomlaiski et al., 2009). For the School version of the scale, Cronbach’s alpha values varied between .91 and .94, which are also similar to the internal consistency values obtained in other studies (DuPaul, Anastopoulos, et al., 1998).
The second purpose of this study was to investigate the factorial structure of the Romanian version of the ADHD-RS-IV in relation to the one proposed by the authors of the scale (DuPaul, Anastopoulos et al., 1998; DuPaul et al., 1997) and confirmed by other studies (Magnússon et al., 1999; Szomlaiski et al., 2009). In addition, the factorial structure of both ADHD-RS-IV versions was verified using CFA for two models: (1) the one-factor model, where each item loads a single factor, and (2) the two-factor model, where odd-numbered items load the IA factor and even-numbered items load the HI factors. For both the ADHD-RS-IV forms in our study, the statistical match indicators showed a better match of our data with the theoretical two-factor model proposed by DuPaul, Anastopoulos, et al. (1998) and DuPaul et al. (1997). A direct comparison between the two models shows that adding a new factor to the initial one-factor model leads to a significant increase in the degree of matching, thus verifying the superiority of the two-factor model. Both versions (Home and School) supported a two-factor conceptualization of ADHD symptoms. The magnitude of items loading on IA and HI factors was greater for the teachers’ data; the patterns of loadings for both data sets also were highly similar, which closely matches the DSM-IV symptom listings. These results are comparable with those obtained in other studies investigating parents’ assessments of ADHD symptoms (Bauermeister et al., 1995). Thus, the two-factor model allows the identification of ADHD clinical subtypes (i.e., predominantly IA or HI, as well as combined IA–HI types) as obtained in DSM-IV clinical studies (Lahey & Carlson, 1991; Lahey et al., 1994).
Another aspect we investigated in the present study was the measurement invariance of the Romanian version of the ADHD-RS-IV Home and School versions, across age, gender, and clinical status of the child, and informant. Both versions showed measurement invariance, which is extremely important when conducting comparisons between different samples. Once measurement invariance is established, we eliminate the possibility that the potential differences between samples may be only measurement artifacts and not actual differences. Therefore, as we obtained strong invariance, we can conclude that there are no differences in the manner in which parents and teachers of girls and boys, younger and older children, as well as those with nonclinical status and children diagnosed with ADHD interpret the ADHD-RS-IV items. Indeed, differences were found in the latent structure of ADHD total symptoms, as well as for the IA and HI dimensions, between boys and girls, younger and older children, and nonclinical samples and clinical samples with ADHD on both parent and teacher ratings.
Finally, normative data from a nationally representative sample of Romanian children were computed to be used in clinical practice and for research purposes in Romania. Through the standardization of the ADHD-RS-IV Home and School versions, we contribute both to research and to clinical practice. Regarding the contribution to research, through this study, (1) we provided interpretation norms for the scores so that a prevalence study on ADHD in Romanian children and adolescents could be conducted, (2) we ensured that differences in the prevalence of ADHD are not accounted for by methodological errors that are dependent on scalar invariance, and (3) we set the stage for cross-cultural international studies. Regarding the clinical implications of this study, through the validation of the Romanian ADHD-RS-IV, we provide Romanian clinical psychologists with a reliable instrument, easy to score and administer, that can be used both in diagnosis and in treatment monitoring. Furthermore, we ensured that in east European countries, clinical psychologists could adhere to the evidence-based assessment guidelines on ADHD and use standardized instruments in their practice.
Several limitations of the present study need to be mentioned. First, given the changes in DSM-5 (APA, 2013) with regard to ADHD symptoms in adolescents, the use of the newest ADHD rating scale, ADHD-5 (DuPaul et al., 2016), would have been more appropriate, as several scores on IA and HI could have been biased in adolescent populations given the formulation of the items. Future studies should investigate cultural influences on the assessment of ADHD and the scale’s longitudinal invariance. Also, even though we assessed teachers’ degree of acquaintance with the child, other variables that were not assessed could have influenced their ratings of the child’s ADHD symptoms (e.g., child performance in school, problematic behavior in the classroom, etc.). The same limitation could be considered in the case of parents, where we did not assess psychopathology (e.g., parental depressive symptoms, parental ADHD).
Conclusions
When assessing children and adolescents suspected of having ADHD, it is important to use reliable and validated behavioral rating scales. If the behavioral assessment scale is developed in one language, it is necessary to verify how the adapted version of the scale functions in a different language. Overall, our findings confirm the validity of the ADHD-RS-IV construct structure in a representative national sample of Romanian children. Even if these results were corroborated in other studies (Zhang et al., 2005), this does not necessarily imply a universal syndrome structure for ADHD-RS-IV. To conclude that an assessment instrument measures the same construct across different societies, it would be necessary to test all components of measurement variance in those individual societies (Ivanova et al., 2007). Our research was limited to Romania, so such an assessment was beyond the scope of this study.
In examining the data, it is also important to consider our findings in relation to the realities of CFA. Support provided by CFA for a particular taxonomic model does not necessarily mean that it is the only model compatible with a particular data set (Hershberger, 1994). Thus, this is a possible limitation of the present research.
In summary, both ADHD-RS-IV versions (Home and School) appear to be suitable instruments for assessing ADHD problems in children, whether in nonclinical research or clinical practice. As with any other rating scale, however, ADHD-RS-IV should not be used as a single instrument for diagnostic purposes. In combination with other assessment procedures (e.g., interviews, direct observation, or psychological testing), however, it appears to make a significant contribution toward increasing diagnostic accuracy, facilitating treatment planning, and assessing treatment outcomes objectively.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Part of this work was supported by a grant of the Romanian Ministry of Research and Innovation, CNCS-UEFISCDI, Project No. PN-III-P4-ID-PCE-2016-0861, Contract No. 146/2017 within PNCDI III, awarded to Dobrean Anca.
