Abstract
Measurement invariance over time (longitudinal invariance) is a core but seldom-tested assumption of many longitudinal studies on adolescent psychosocial development. In this study, we evaluated the longitudinal invariance of a brief measure of adolescent mental health: the Social Behavior Questionnaire (SBQ). The SBQ was administered to participants of the Zurich Project on the Social Development of Children and Youths in up to four waves spanning ages 11 to 17. Using a confirmatory factor analysis approach, metric invariance held for all constructs, but there were some violations of scalar and strict invariance. Overall, intercepts tended to increase over time while residual variances decreased. This suggests that participants may become more willing or able to identify and report on certain behaviors over time. The noninvariance was not practically significant in magnitude, except for the Anxiety dimension where artifactual increases over development would be liable to occur if invariance is not appropriately modeled. Overall, results support the utility of the SBQ as an omnibus measure of psychosocial health across adolescence.
A key goal of longitudinal research in adolescence is to illuminate processes of development. Usually this is achieved through the examination of stability and change in emotional, psychological, and behavioral traits over several years. The validity of conclusions drawn from such research relies on the availability of comparable trait estimates over development. This in turn requires at least partial longitudinal invariance, namely, that a subset of items capture the same construct on the same measurement scale over time (Edwards & Wirth, 2012; Widaman, Ferrer, & Conger, 2010). Measurement invariance for items can be violated for a variety of reasons, especially in periods of substantial developmental change such as adolescence. When this happens and it is not appropriately modeled, apparent changes over time could be due to changes in the way a construct is measured, true changes could be masked, or an entirely different construct could be measured at different time points (e.g., Edwards & Wirth, 2009).
Core constructs in adolescent psychosocial development research include conduct issues, depression, anxiety, attention deficit hyperactivity disorder (ADHD), and prosocial behavior. A sizeable body of research has sought to characterize development of these dimensions in adolescence. This includes studies of average developmental trajectories, developmental trajectory subtypes, developmental interrelations, and studies identifying predictors and outcomes of developmental trajectories (e.g., Carlo, Crockett, Randall, & Roesch, 2007; Crocetti, Klimstra, Keijsers, Hale, & Meeus, 2009; Dekker et al., 2007; Luengo Kanacri, Pastorelli, Eisenberg, Zuffianò, & Caprara, 2013; Marmorstein, 2009; Martino, Ellickson, Klein, McCaffrey, & Edelen, 2008; Murray, Eisner, & Ribeaud, 2016b; Murray, Obsuth, Eisner, & Ribeaud, 2017; Murray, Obsuth, Zirk-Sadowski, Ribeaud, & Eisner, 2016; Nantel-Vivier et al., 2009; Van Oort, Greaves-Lord, Verhulst, Ormel, & Huizink, 2009). Studies of this kind have, for example, suggested that across adolescence there are general decreases in antisocial behavior, ADHD symptoms, anxiety, depression, and prosociality, with a possible rebound in late adolescence for the latter. However, these average trends are in the context of considerable variation across individuals with, for example, growth mixture analyses, often revealing a nontrivial subgroup for whom levels of these dimensions increase across development (e.g., Crocetti et al., 2009). Studies in this research area have also delineated pathways by which these dimensions may influence one another. Conduct problems may, for example, put an individual at risk of academic and social difficulties that in turn increase the risk of anxiety and depression (e.g., van Lier et al., 2012).
The above-mentioned studies and other longitudinal studies of adolescent psychosocial development rely on trait estimates that can be validly compared across development; however, the array and pace of changes that occur during adolescence make this a challenge. Over development, constructs may change in nature, particular indicators may lose their developmental appropriateness, or samples may reach a floor or ceiling on items that previously discriminated well between different trait levels (e.g., Edwards & Wirth, 2012). Early adolescence, for example, sees a spike in certain types of antisocial behavior such as rule-breaking and noncompliance with authority. For most individuals, these behaviors decline in late adolescence and adulthood (e.g., Barker et al., 2007). As such, endorsing an item measuring one of these behaviors would tend to index a greater degree of underlying severity in late adolescence than the same behavior endorsed in early adolescence. Similarly, the attention deficit versus hyperactivity/impulsivity symptoms of ADHD have previously shown evidence of differential developmental trajectories. The former has shown a potential curvilinear trajectory while the latter show a more definite and steady decline over development (e.g., Murray, Obsuth, et al., 2016). Over the course of adolescence this could, for example, lead to increasing bifurcation of a previously unitary ADHD construct into separate inattention and hyperactivity/impulsivity constructs. More generally, items referring to symptoms or behaviors within the context of typical early adolescent social environments and activities (e.g., school) may be less relevant by late adolescence.
When items cannot be considered comparable across time, valid inferences about changes in the underlying trait rely on identifying and modeling the nature and extent of noncomparability. “Comparability” comes in different degrees and types that can be modeled and tested statistically. A useful framework for doing so is that of longitudinal measurement invariance within confirmatory factor analysis (CFA; e.g., see Millsap & Cham, 2012). Longitudinal measurement invariance is when expected observed score distributions given trait levels are independent of the wave at which the measure was administered. Longitudinal invariance using CFA tests a slightly weaker version of this, namely, that the expected mean and variance of observed score distributions given latent trait levels are independent of measurement wave. In this framework, a measure may show (from weakest to strongest) no invariance, configural invariance, metric invariance, scalar invariance, or residual invariance.
Configural invariance means that only the pattern of factor loadings for a measure is the same across time. Metric invariance is where both factor loading patterns and factor loading magnitudes are equal across time. When metric invariance holds, comparisons of factor variances and covariances are supported, for example, in cross-lagged and autoregressive panel models. When metric invariance does not hold, an inventory may not be measuring the same construct across time. Scalar invariance is when factor patterns, loadings, and intercepts are equal across time. When scalar invariance holds, inferences about factor mean differences over time are supported. This is relevant when, for example, fitting second-order growth curve models. Finally, residual (or strict) invariance is when factor patterns, loadings, intercepts, and item residual variances are equal over time. When this holds, differences in means and variances of the observed scores can be attributed to differences on the underlying latent factors. Only in this case are inferences about stability and change in a trait over time based on observed scores (e.g., sum scores) supported. Where any level of the above-described invariance assu mptions are not met, it is often possible to nonetheless obtain valid comparisons of latent constructs across time. To do so requires that at least some items are longitudinally invariant, giving “partial invariance.” Partial invariance often suffices provided that the noninvariance in the remaining items is explicitly modeled (e.g., Edwards & Wirth, 2012).
Surprisingly few studies of adolescent development report longitudinal invariance analyses either in their own right or to support other longitudinal analyses. Some studies have suggested that high levels of invariance across adolescence can be achieved for specific measures of mental health dimensions, such as ADHD, depression, anxiety, and conduct problems (e.g., Leopold et al., 2016; Motl, Dishman, Birnbaum, & Lytle, 2005; Sterba et al., 2010; Verhoeven, Sawyer, & Spence, 2013) while others have identified noninvariance (Mathyssek et al., 2013; Sterba et al., 2010). In this study, we sought to build on this limited evidence base and evaluate longitudinal invariance in a brief measure of five dimensions of psychosocial functioning across development: the Social Behavior Questionnaire (SBQ; Tremblay et al., 1991).
The SBQ shares origins with the Strengths and Difficulties Questionnaire (Goodman, 1997). Its first reported use was in a study by Tremblay et al. (1991) where two preexisting scales were combined: 28 items from the Preschool Behavior Questionnaire (Behar & Stringfield, 1974), itself an adaptation of the Children’s Behavior Questionnaire (Rutter, 1967), and 10 items from the Prosocial Behavior Questionnaire (Weir & Duveen, 1981). Though originally developed for young children, it has since been applied in older children and adolescent samples. It has been adopted in several large-scale studies internationally, where variants have been administered in English, French, and German. For example, items of the SBQ were integrated in the behavior evaluations in the Canadian National Longitudinal Survey of Children and Youth (Sprott, Jenkins, & Doob, 2000), the Nuremberg-Erlangen Prevention and Development Study (e.g., Lösel & Stemmler, 2012), the Zurich Project on the Social Development of Children and Youths (z-proso; Eisner & Ribeaud, 2007), and the Quebec Longitudinal Study of Kindergarten Children (QLSKC; Rouquette et al., 2014). It is used in self, parent, and teacher report form.
In spite of the large number of published studies using the SBQ items, there have been only a small number of previous dedicated studies of its psychometric properties. The few that have been conducted have generally supported the factorial validity, criterion validity, and reliability of the SBQ items (e.g., Murray, Eisner, & Ribeaud, 2017; Murray, Eisner, & Ribeaud, 2016a; Tremblay et al., 1991; Tremblay, Vitaro, Gagnon, Piché, & Royer, 1992). Murray et al. (2017), for example, analyzed the ranges of measurement for which the teacher-reported SBQ items reliably captured the dimensions of prosociality, externalizing, ADHD, and internalizing in z-proso at ages 7 through to 15. The latter three traits were operationalized using bifactor measurement models where the specific dimensions of ADHD were inattention and hyperactivity/impulsivity, the specific dimensions of internalizing were anxiety and depression, and the specific dimensions of externalizing were physical aggression, oppositional defiant disorder/conduct disorder, and reactive aggression. They found that SBQ items could reliably capture a wide range of trait levels for the general factors across ages 7 to 15 years. Murray, Obsuth, et al. (2017) reported a factor analysis of all self-reported SBQ items together with the newly developed Violent Ideations Scale (Murray et al., 2016a) in the z-proso sample when the participants were aged 17 years. All items from the SBQ loaded on the intended factors (prosociality, internalizing, ADHD, reactive aggression, and proactive aggression). In contrast to the SBQ reliability study by Murray et al. (2017), this study did not support a distinction between anxiety and depression; however, Murray, Obsuth, et al. (2017) cited concerns about overfactoring, and it is possible that extracting further factors would have yielded separate anxiety and depression factors. Arguably, understanding the longitudinal measurement properties of an instrument is an important step in its validation and in construct validation more broadly (e.g., Edwards & Wirth, 2009); however, this has not been tested for any version of the SBQ across any developmental period. In this study we, therefore, evaluate longitudinal measurement invariance of the SBQ in the z-proso sample.
Method
Participants
Data came from the z-proso study. This is a longitudinal cohort study of youth, focusing on the development of prosocial and antisocial behaviors. The study began in 2004 when the participants were aged 7 years and entering school. Participants were selected according to a school-level stratified random sampling procedure that took into account school size and location. All children entering the first grade at selected schools in that year were invited to participate. The study included separate child and parent intervention components in the early waves; however, as there was little evidence that they had any substantive short- or long-term effects, it is usually judged reasonable to treat the data as observational (e.g., Averdijk, Zirk-Sadowski, Ribeaud, & Eisner, 2016; Malti, Ribeaud, & Eisner, 2011). Comprehensive accounts of recruitment, participant characteristics, attrition, and assessment procedures can be found in previous publications (e.g., Eisner & Ribeaud, 2007) and on the z-proso website (http://www.cru.ethz.ch/en/projects/z-proso.html). The current study focusses on the latter four measurement waves when the majority of participants were aged 11, 13, 15, and 17 years, respectively. Across these waves, 1,523 (51% male) participants contributed data, representing 91% of the original target sample.
Measures: Prosociality, Anxiety, Depression, ADHD, and Aggression
Prosociality, anxiety, depression, ADHD, and Aggression were measured with the SBQ. Z-proso includes some adaptations of the SBQ, mainly additional aggression items reflecting the fact that a major research theme of the study is antisocial behavior development. In addition, whereas the original SBQ was on a 3-point response scale, the version administered in z-proso offers respondents a 5-point scale. In our longitudinal invariance analyses, we focus on items that were common across all measurement waves to simplify the interpretation of any noninvariance identified. The majority of these items were part of the original SBQ reported in Tremblay et al. (1991). One anxiety item, two depression items, and two aggression items were added in z-proso to maintain developmental appropriateness in the adolescent period. Overall, the SBQ administered in z-proso included eight items measuring prosociality that were common across all measurement waves (age 11, 13, 15, and 17), four measuring depression, and four measuring anxiety. In addition, four items measuring ADHD and eight measuring aggression (four each for the subtypes of proactive and reactive aggression) were administered at the latter three measurement waves (age 13, 15, and 17) only. Thus, we analyze a slightly different set of items as compared to the original SBQ items developed for preschoolers. Abbreviated item contents are provided in the Results tables. All items are measured on a 5-point Likert-type scale from never to very often. Apart from the anxiety and depression items, which referred to the frequency of behavior/symptoms in the past month, all items referred to frequency of behavior/symptoms in the past year. All were administered in paper-and-pencil format in German: the official language of the study location.
Statistical Procedure: Longitudinal Factorial Invariance
Longitudinal factorial invariance was assessed within a CFA framework. We treated items as continuous because it is usually reasonable to treat ordered-categorical items with at least five response options as continuous (e.g., Rhemtulla, Brosseau-Liard, & Savalei, 2012). Benefits of doing so include better accounting for missingness (because full information maximum likelihood estimation can be employed), a larger simulation evidence base to draw on to guide model selection based on model fits (e.g., Chen, 2007; Cheung & Rensvold, 2002; Meade, Johnson, & Braddy, 2008), and the availability of information theoretic criteria to further guide model selection (e.g., Raftery, 1995). Furthermore, Sass, Schmitt, and Marsh (2014) presented simulation study results suggesting that fit comparisons for invariance testing using weighted least squares means and variances estimation do not perform well.
Comparisons across a series of increasingly constrained models provided information on the level of invariance that could be achieved for each trait. We report χ2 difference tests for information but note that these are likely to be overly sensitive to minor misspecifications given our large sample size (e.g., Meade et al., 2008). As such, for model selection, we relied primarily on the approximate fit indexes of comparative fit index (CFI), Tucker–Lewis index (TLI), and root mean square error of approximation (RMSEA), which are less influenced by sample size and model complexity than the χ2 difference test. In particular, we use the cutoff criteria defined by Chen (2007), developed using a simulation study of different types of invariance at different sample sizes. According to these criteria, metric invariance would hold in our sample when the CFI decreased by less than .010, RMSEA increased by less than .015, and standardized root mean square residual (SRMR) increased by less than .030 with the addition of metric invariance constraints. Scalar invariance would hold if CFI decreased by less than .010, RMSEA increased by less than .015, and SRMR increased by less than .010 with the addition of scalar invariance constraints. Residual (or strict) invariance would hold if CFI decreased by less than .010, RMSEA increased by less than .015, and SRMR increased by less than .010 with the addition of scalar invariance constraints. Noting that these criteria did not have power to detect minor violations of invariance, Meade et al. (2008) suggest an increase in CFI of >.002 should be used to detect noninvariance; however, given the large number of comparisons to be conducted overall, we elected to use Chen’s (2007) more conservative criteria to avoid the detection of a large number of trivially small violations of invariance.
To test configural, metric, scalar, and residual invariance, the following series of increasingly strict models were fit to test invariance. First, in the configural model, patterns of loadings were fixed equal over time but loadings and intercepts were free to vary. Scaling and identification were achieved by fixing the mean and variance of the latent factor at baseline to 0 and 1, respectively, and by fixing the loading and threshold of a reference item to equality over all time points. This method assumes that the reference item is invariant over time. Choosing a reference item that is not invariant can lead to the appearance of noninvariance in truly invariant items and/or to the appearance of invariance in truly noninvariant items (e.g., Yoon & Millsap, 2007). However, there was neither prior empirical evidence nor strong theoretical rationale on which to base the selection of reference variable for the SBQ. As such, we provisionally selected the first item in each subscale to be the reference variable and then checked that the invariance constraints with these items were not associated with large modification indices at either the metric or scalar invariance model stages. If they were, we allowed these constraints to be freed and instead thereafter relied on other items to act as reference variables. Residual covariances between the same items measured over time were also estimated.
Configural invariance was judged to hold when the configural model had RMSEA < .08, SRMR < .08, and TLI and CFI > .95 (e.g., Hu & Bentler, 1999; Schermelleh-Engel, Moosbrugger, & Muller, 2003). In the metric model, equality constraints over time were added on the remaining first-order loadings. In the scalar model, equality constraints over time were added on the remaining intercepts. In the strict invariance model, residual variances were fixed equal over time. If at any stage invariance was not supported, modification indices and expected parameter changes were used to identify specific constraints that did not hold. These were iteratively released, retesting invariance with the release of each individual constraint until either partial invariance held or there were only two items left with invariance constraints at that level. In this latter case, invariance was judged not to hold at that level. No further constraints were added to items that did not show invariance at a lower level, for example, scalar constraints were not placed on items that did not show metric invariance.
Finally, the practical consequences of any noninvariance identified were investigated. To do this, we compared two estimates. The first was a linear slope factor mean from a growth curve models fit using a model that assumed full invariance (all loadings, intercepts, and residual variance assumed equal over time). The second was the same parameter from a model that assumes the highest level of invariance that was actually attained. The difference provides a quantification of the bias that could be expected to result from the noninvariance, if not appropriately modeled.
Results
Overview
Invariance was not supported at any level according to the χ2 difference test; however, as discussed above, this test is likely not appropriate for our large sample size as it will identify even trivially small misspecifications as statistically significant. Partially invariant models were achieved in all cases according to the fit criteria of Chen (2007). The practical significance of the noninvariance identified was overall small: Unstandardized slope estimates from partially invariant versus (falsely assumed) fully invariant models generally differed only at the second decimal place. The one exception was anxiety, for which slopes were overestimated by more than 50% when longitudinal invariance was incorrectly assumed.
Prosociality
Fits for the prosociality models are provided in Table 1. The configural model was a single-factor model. Both configural and metric invariance held, but scalar invariance did not. Iterative release of equality constraints on intercepts in the following sequence resulted in a partially invariant model (M2a in Table 2): Item 48 at age 11, Item 47 at age 17, Item 41 at age 17, and Item 48 at age 13. Adding residual invariance constraints, fit statistics suggested further noninvariance. Based on modification indices and expected parameter changes, constraints on the item residual variances for Item 47 at age 11 and then Item 41 at age 11 were released. At this point, partial residual invariance was judged to hold. Parameters for this final model (M3a) are provided in Table 2. These show that for Item 48, the intercept increased over time. In addition, the residual variances for Items 41 and 47 were larger at age 11 than at subsequent measurement points. The linear slope factor mean from a latent growth curve model fit using the final measurement model developed in these analyses was −0.290. The model assuming full measurement invariance across development yielded a corresponding estimate of −0.209.
Model Fits for Prosociality Invariance Models.
Note. df = degrees of freedom; CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion. Final model indicated in boldface.
Parameters for Most Invariant Prosociality Model.
Note. Noninvariant parameters are indicated in boldface.
Anxiety
Fits for the anxiety models are provided in Table 3. The configural model was a single-factor model. Both configural and metric invariance held. Scalar invariance did not hold, but partial invariance was achieved with the iterative release of scalar constraints on Item 1 measured at age 11 and then the same item measured at age 13. No further noninvariance was identified with the addition of residual invariance constraints on the scalar invariant items. Partial residual invariance was, therefore, judged to hold. Parameters for this model are provided in Table 4. These show that the intercept for Item 4 increased across ages 11, 13, and 15. The linear slope factor mean from a latent growth curve model fit using the final measurement model developed in these analyses was 0.257. The model assuming full measurement invariance across development yielded a corresponding estimate of 0.544.
Model Fits for Anxiety Invariance Models.
Note. df = degrees of freedom; CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion. Final model indicated in boldface.
Parameters for Most Invariant Anxiety Model.
Note. Noninvariant parameters are indicated in boldface.
Depression
Fits for the depression models are provided in Table 5. The configural model was a single-factor model and showed good fit. Both configural and metric invariance held, but scalar invariance did not. The scalar invariance constraint on Item 63 at age 13 was removed to achieve partial scalar invariance. Partial residual invariance was achieved with the addition of residual invariance constraints to all remaining items. The parameters from this model are provided in Table 6. Item 4 had a larger intercept at age 13. The linear slope factor mean from a latent growth curve model fit using the final measurement model developed in these analyses was 0.716. The model assuming full measurement invariance across development yielded a corresponding estimate of 0.725.
Model Fits for Depression Invariance Models.
Note. df = degrees of freedom; CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion. Final model indicated in boldface.
Parameters for Most Invariant Depression Model.
Note. Noninvariant parameter is indicated in boldface.
Attention Deficit Hyperactivity Disorder
Fits for the ADHD models are provided in Table 7. The configural model for ADHD was a single-factor model. Configural, metric, scalar, and residual invariance all held. Parameter estimates from the residual invariance model are provided in Table 8. The slope factor mean for a linear growth curve model using this fully invariant ADHD model was 0.114.
Model Fits for Attention Deficit Hyperactivity Disorder Invariance Models.
Note. df = degrees of freedom; CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion. Final model indicated in boldface.
Parameters for Most Invariant Attention Deficit Hyperactivity Disorder Model.
Aggression
Fits for the aggression models are provided in Table 9. The configural model was a first-order oblique model with first-order factors: proactive aggression and reactive aggression. The configural model showed reasonable fit, and configural invariance was judged to hold. Metric but not scalar invariance held. Release of the intercept constraint on Item 61 at age 17 was necessary to achieve partial scalar invariance. To achieve partial residual invariance, it was necessary to release the constraints on the residual variance of Item 37 at age 17 and then on Item 72 at age 13. Parameter estimates for this model are provided in Table 10. These show that the intercept for Item 61 was lower at age 17 while the residual variance for Items 37 and 72 was larger at earlier waves. The linear slope factor mean from a latent growth curve model fit using the final measurement model developed for reactive aggression was −0.041 and for proactive aggression was −0.210. The models assuming full measurement invariance across development for these constructs yielded corresponding estimates of −0.084 for reactive aggression and −0.202 for proactive aggression.
Model Fits for Aggression Invariance Models.
Note. df = degrees of freedom; CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion. Final model indicated in boldface.
Parameters for Most Invariant Aggression Model.
Note. Noninvariant parameters are indicated in boldface.
Discussion
In the current study, we examined the important but seldom-tested assumption of longitudinal factorial invariance for five dimensions of adolescent psychosocial functioning measured by the SBQ. We found that items largely functioned equivalently across waves, supporting their use in longitudinal analyses. Where noninvariance was identified, the most consistent pattern was larger residual variances at earlier time points. This suggests that the adolescents may have become better able to report on the presence of certain behaviors or symptoms over time. The potential bias arising from falsely assuming invariance with these data (e.g., by using sum scores) is not likely to be substantial, except for developmental analyses involving anxiety.
Longitudinal invariance analyses are rarely reported in studies of adolescent development. This is in spite of the fact that the majority of conclusions drawn about psychosocial development from longitudinal data rely on the assumption that at least metric invariance holds. Furthermore, the selection of measures for longitudinal studies seldom explicitly considers the degree of comparability of measures over the relevant phases of development as a criterion. In this study, we use this evaluated longitudinal invariance for five dimensions of adolescent mental health measured by the SBQ. All subscales showed at least metric invariance which is consistent with the idea that they measure the same constructs over adolescence. This makes the SBQ a good candidate for use as a brief omnibus measure of adolescent psychosocial development.
Scalar invariance was violated for some items across the SBQ subscales. There was a general tendency here for item intercepts to increase over measurement waves. As such, for a given latent trait level, expected item scores were higher at later time points. The implication of this is that a researcher using observed scores (e.g., sum score for a subscale) would see an artifactual increase (or an attenuated decrease) in levels of the traits measured with these items over adolescence. However, as so few items were affected, it is likely that unbiased comparisons of levels over time could be made using a partial measurement invariance model in which the scalar invariance constraints that did not hold were not imposed (Byrne, Shavelson, & Muthén, 1989; see also Ferrer, Balluerka, & Widaman, 2008). Comparisons of latent growth curve models from models appropriately modeling noninvariance versus models that assumed full invariance revealed substantial discrepancies only for anxiety. For anxiety, the positive slope over time was estimated as 52% larger when assuming invariance because the model attributed an increase in an item intercept to an increase in factor means across time.
A possible explanation for the trend toward increasing intercepts over time is that adolescents become more attuned to the behaviors and symptoms about which the items ask over time. This could be a function of the repeated administration of the questionnaires whereby participants are cumulatively primed to detect certain symptoms and behaviors. It could also be a function of increasing capacity to identify and report on these same symptoms and behaviors arising from maturity. An alternative explanation would be that participants become more comfortable disclosing information about their negative behaviors and symptoms over time as they build trust with the study over measurement waves; however, the fact that the same trend was seen in prosociality—a trait with positive connotations—calls this interpretation into question. The one item that deviated from this pattern was an item measuring tendency to experience boredom, an indicator of depression. This item had a larger intercept at age 13 than at other ages. One possibility is that this reflects a normative elevation of boredom specifically around this age (Spaeth, Weichold, & Silbereisen, 2015).
There were also some violations of residual invariance. In general, residual variances were larger at age 11, suggesting that measurement error decreased over time. This is likely to reflect increased familiarity with the questionnaire on repeated administrations and an increase in the reliability of self-reports that comes with maturity. However, the fact that these violations were few and generally of a small magnitude suggests that by age 11 self-reports are not substantially less reliable than those at older ages.
There are implications of these findings both for users of the z-proso data set specifically and researchers of adolescent development more generally. Users of the data set should ideally adopt the measurement models for the anxiety, depression, prosociality, and aggression outlined in the Results section, especially for anxiety. For certain models with a high degree of computational complexity, such as those involving a large number of latent interactions, using factor scores estimated from the measurement models would represent a more practical solution. In these cases, the researcher should confirm that the determinacy of the factor scores is adequate, for example, >.90 (Gorsuch, 1983). For ADHD, which showed full invariance over development, using sum scores over latent variable measurement models would not be expected to introduce substantial bias. However, there would nonetheless be other benefits to using a latent variable measurement model, such as greater power due to the disattenuation for unreliability and the ability to test model fit and identify misspecifications.
While these results provide validity support for the SBQ in showing that most items remain comparable across development, users of the SBQ in other studies should independently investigate invariance and develop appropriate measurement models to take account of any violations identified. There are no guarantees that invariance results would generalize across studies because the methodological differences between studies using the SBQ and the different languages and cultural contexts in which it is administered could influence longitudinal invariance. For example, the SBQ is administered on a 3- rather than 5-point response format in some studies. Comparisons of invariance results across different studies using the SBQ would help home in on the features generating the noninvariance.
More generally, our study illustrates that in spite of the myriad changes occurring across adolescence, it is possible to construct measures that capture dimensions such as anxiety, depression, ADHD, aggression, and prosociality in a comparable way across development. This is critical to robust research into developmental trends in these dimensions. Any study seeking to make inferences about changes in variances, covariances, and means (or statistics derived from these) should follow a procedure similar to that outlined in our Method and Results section, in order to build appropriate measurement models to account for noninvariance prior to testing their substantive hypotheses. In some cases, it may be necessary to simultaneously investigate invariance with respect to other variables. For example, tests of differential subgroup developmental trajectories (e.g., sex differences in development) would call for investigations of invariance by both subgroup and measurement wave.
The fact that the SBQ showed high levels of invariance over time likely reflects the fact that the wording of items is relatively general and not tied to specific contexts (e.g., school) or activities (e.g., schoolwork). When planning longitudinal studies, including a core set of “context-free” or “developmentally neutral” items that could be expected to show invariance over development may be crucial to ultimately obtaining comparable trait estimates over longer time spans. This is not to discourage the inclusion of items that are specific to some developmental periods; omitting these may result in crucial manifestations of traits being missed. For these items to be included in studies of change over time, however, they need to be anchored via developmentally invariant items (e.g., Edwards & Wirth, 2009). The SBQ items may, for example, serve as useful anchors when developing new measures for longitudinal studies or administered alongside more comprehensive measures of the dimensions it measures.
Finally, it is important to consider the limitations of the approach of the current article. First, factorial invariance does not guarantee that a measure captures the same psychological phenomenon over time; it is possible for metric invariance to hold, for example, when the psychological process underlying responses is changing (e.g., Widaman, Little, Geary, & Cormier, 1992). Second, we selected fit criteria for detecting invariance to protect against making Type 1 errors. However, this also increases the likelihood of missing minor violations of invariance. Third, we focused on items that were identical over multiple waves. As such, the measures may not have been representative of the traits during particular time periods. Fourth, although overall tests of invariance at the scale level tend to be relatively robust, this is not necessarily true at the item level where, for example, iterative release of constraints based on modification indices or the selection of a noninvariant reference variable can result in misidentifying the location of noninvariance (e.g., Johnson, Meade, & DuVernet, 2009). As such, our item-level inferences should be regarded with more caution. Finally, we did not have any a priori hypotheses regarding which items should be noninvariant and in what way (e.g., whether intercepts should increase over development). As such, our interpretations of the noninvariance were entirely post hoc. Future studies may benefit from deriving specific a priori hypotheses about noninvariance across time from developmental theories of the traits being analyzed.
Conclusions
Although adolescence is a time of considerable change, it is possible to identify items that show comparable measurement properties across this period: a prerequisite for supporting valid inferences about development. Where measurement invariance was violated, it tended to be the case that residual variances were larger at earlier waves while intercepts were larger at later waves. In the SBQ, only the anxiety subscale showed practically significant levels of noninvariance.
Footnotes
Acknowledgements
We are grateful to the children, parents, and teachers who provided data for the z-proso study and the research assistants involved in its collection.
Authors’ Note
Ingrid Obsuth is now affiliated to University of Edinburgh, Edinburgh, UK and Denis Ribeaud is now affiliated to University of Zurich, Zurich, Switzerland.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research is supported by the Jacobs Foundation (Grant 2010-888) and the Swiss National Science Foundation (Grants 100013_116829 & 100014_132124) is also gratefully acknowledged.
