Abstract
Historically, the assessment of longitudinal construct validity in the field of psychosocial measurement involved defining hypotheses and calculating correlation coefficients using scores based on 2 measures at 2 or more time points. In the context of patient-reported outcomes, this evolved into sensitivity to change and responsiveness, including the computation of effect size estimates of change, standardized response means, and indices such as Guyatt’s statistic. Cross-sectional analyses or analyses based on 2 time points have been the standard practice. Evolving conceptualizations have incorporated more than 2 time points and have included depictions of individual trajectories of change in multiple measures, structural equation models, and mixed modeling techniques. The focus of this article is on methods to evaluate longitudinal construct validity. We describe a sample of these methods and provide considerations and recommendations for designing a thoughtful longitudinal construct validity evaluation of clinical outcome assessments.
Keywords
Introduction
While the traditional textbook definition of validity is “the extent to which an assessment measures what it is supposed to measure,” 1 –5 a comprehensive definition of validity refers to the extent that the proposed uses and interpretations of assessment scores are supported by theory and evidence. 6 –13 Validity is not simply yes or no—it is whether a particular concept is adequately measured by an assessment and, more importantly, whether the scores derived from the assessment can support inferences and decisions that are expected to be made using the assessment. As stated in the FDA’s guidance on patient-reported outcome (PRO) measures, “the adequacy of an instrument’s development and testing is specific to its intended application in terms of population, condition, and other aspects of the measurement context for which the instrument was developed.” 14 Indeed, a medical product labeling claim based on a clinical outcome assessment (COA) must be “consistent with the instrument’s documented measurement capability … to ensure the conclusions drawn using instrument scores are valid.” 14
It has also long been stated that “validation is never finished.” 15 It is an “unending” 12 or at least an “ongoing” 16,17 and “continuing” process. 10 What is intended by these assertions is that validity cannot be definitively established or proven—much like a hypothesis cannot be proven—and an assessment is never conclusively validated. Instead, analyses provide and accumulate evidence for the validity and interpretation of the scores in a particular setting and within a given patient population or context of use.
For example, COAs used to support endpoints in clinical trials or observational studies that are considered to validly measure the physical functioning of the general adult population in the United States are not necessarily valid to evaluate changes in the physical functioning of elderly heart failure patients in the context of a multicenter international clinical trial. Nor are those COA scores automatically valid measures for comparing postsurgical differences in physical functioning between groups of patients receiving different types of knee implants. As another example, COA scores used by clinicians in practice for the purpose of tracking the worsening of physical symptoms in patients with a particular neuromuscular disorder will not necessarily support a claim of treatment benefit in medical product labeling or even symptom improvement in a clinical trial.
The validity of an assessment can be divided into numerous types, including face validity, content validity, criterion validity, and construct validity. 12,18,19 In the United States, regulatory review of COA scores proposed as key endpoints is often led by the Food and Drug Administration (FDA) COA staff (formerly Study Endpoints and Labeling Development [SEALD]). Guidance by this team has focused primarily on content validity and construct validity: Content validity as supported through qualitative input from patients during the development of a COA to provide evidence that the measure captures the aspects that are relevant to capture; construct validity as primarily supported through quantitative results that COA scores are consistent with a priori hypotheses (including hypotheses describing the internal COA relationships among item scores and external relationships among items and scale scores with scores from other related and unrelated measures). 7,12,14,19 Although the FDA guidance specifies PROs in its title, the methods and principles of the PRO guidance are to be applied to the evaluation of all COAs, including observer-reported outcomes (ObsROs) and clinician-reported outcomes (ClinROs).
The focus of this article is on construct validity, specifically the extension of the evaluation of construct validity beyond a cross-sectional assessment to incorporate data from multiple time points: a longitudinal assessment of construct validity. Numerous statistical techniques have been developed to simultaneously model multiple outcomes collected across multiple time points in the study of change and factors that impact change. 20 –26 Our overall objective is to describe methods, including a sample of emerging techniques, used for evaluating the construct validity of COAs over time with considerations and recommendations for designing a thoughtful COA longitudinal construct validity evaluation.
Construct Validity Methods
As there are multiple ways to classify aspects of validity, there are multiple ways to classify construct validity. For the purpose of this paper, any psychometric assessment that begins with a hypothesis about how COA item or scale scores behave is construct validity. Using this definition, evidence for a COA’s construct validity can be gained through the use of three key types of analyses: correlational, known groups, and responsiveness analyses.
While correlational analyses are considered the primary method, we have included known groups and responsiveness within construct validity because results from these methods can provide important information about relationships among scores. Again, the essential requirement for the application of these evaluations is the development of hypotheses based on the constructs intended to be measured by the target COA and the constructs or groups defined by the available external measures or criteria. We contend that results reported without this theoretical step are not evaluations of construct validity.
Correlational Analyses
In all construct validity applications, hypotheses are proposed regarding the relationships between the constructs intended for measurement by the COA and available external measures. These hypotheses should indicate the direction (positive or negative) of the construct relationships and the magnitude (either in strength or relative to other planned correlations). Typically, correlation coefficients are computed between the COA and external measures to assess the strength and direction of the hypothesized associations. Correlational analyses treat the 2 outcome measures symmetrically. Other regression-based methods to test associations require the COA to be predicted by an external measure, and there are several excellent references describing the extension of these methods in both cross-sectional and longitudinal frameworks. 1,27 –29 However, we often compute correlation coefficients and employ Cohen’s 30 correlation rule of thumb to describe the strength of expected correlations (strong: r ≥ 0.50; moderate: 0.10 ≥ r < 0.50; weak: r < 0.10), but in some cases we may be able to hypothesize only in terms of relative strength.
An optimal correlational analysis will include convergent (similar) and divergent (dissimilar) construct validity hypotheses, with internal hypotheses based on the item scores of the COA and external hypotheses constructed for the COA scores with scores from other measures. The pattern can be formalized and tested as in the context of a multitrait-multimethod matrix. 12,31 Alternatively, Zou 32 provides an illustration of statistically testing individual correlation values against a prespecified value (eg, H0: ρ = 0.5) rather than relying on the traditional null (H0: ρ = 0). Cross-sectional or longitudinal construct validity is supported when the observed correlation values are as hypothesized (in terms of magnitude, sign, and pattern).
Reeve and colleagues report the results of a survey of members of the International Society for Quality of Life Research (ISOQOL) that sought to establish minimum standards for PRO measures. 5 Fifty-five percent would require that a measure have evidence supporting its construct validity, including “documentation of empirical findings that support predefined hypotheses on the expected associations among measures similar or dissimilar to the measured PRO”; 44% considered this “desirable but not required.” 5
Cross-sectional
After hypotheses have been finalized, the computations can be as simple as a series of Pearson or Spearman correlation coefficients between the target COA and another measure that can be computed using data from a single time point or a series of single time points. These snapshots provide evidence (or lack of evidence) in support of the hypothesized relationships. Comparisons cannot be statistically tested for independent correlation coefficients across time and measures unless modeled simultaneously.
Longitudinal
We can capitalize on the nature of longitudinal COA data for a statistically testable evaluation of construct validity over time. The simplest case is the computation of a correlation using change scores from 2 time points; for increased complexity, change scores across more than 2 time points can be evaluated (Figure 1 shows individual trajectories of several individuals for a single COA).

Individual clinical outcome assessment score trajectories with group mean (dashed line).
With this approach, hypotheses can account for the time-varying nature of the relationships, thereby providing a context for testing differences in the relationships over time. For example, if an established treatment is included in the study, it may be hypothesized that the relationship between a COA and a related clinical variable’s change will be more strongly associated after several weeks of treatment than during the study run-in period.
There are numerous approaches and software programs for simultaneously modeling multiple dependent variables (in our application, the target COA and external measures). Mixed models with both fixed- and random-effects structures present one mathematical framework to estimate the correlations jointly. Results from these models can provide cross-sectional and longitudinal information, including statistical comparisons, supporting the construct validity of the target COA.
One approach presented in the statistics literature is to fit the generalized linear mixed model (GLMM) to accommodate both continuous and discrete dependent variables. 20,33,34 Faes and colleagues 35 describe 2 ways to jointly model and test the correlation between 2 dependent variables, which can be implemented in SAS PROC GLIMMIX: (1) in the residual error term using an unstructured correlation matrix (ie, each pair of outcomes has its own correlation coefficient) 35,36 or (2) through a shared random effect.
Another approach used extensively in the social sciences is to simultaneously fit these curves (Figure 1) in the structural equation modeling (SEM) framework. 23,24,37 –39 Within the SEM framework, these models are typically referred to as latent growth models (LGMs). Essentially, LGMs are a special case or mapping of the linear (or the generalized linear) mixed models and provide an estimate of the between-person differences in within-person change. 40 Dependent LGMs can be implemented in SEM software such as MPlus. (See Box 1.)
In growth-mixture models, the fixed effect represents the mean trajectory pooled across all persons in the sample, and the random effect represents the variance of the individual trajectories around the mean trajectory. In the LGM framework, growth parameters are modeled as a set of latent constructs. This requires that each loading of the observed variables (eg, the time-varying COA observations) and latent growth parameters (eg, intercept and slope) be fixed to a particular value. Typically, the loadings on the latent slopes are fixed to polynomial regression coefficients (eg, linear, quadratic, cubic), and the loadings on the latent intercepts are fixed to 1.
The following equations present an example parameterization of an unconditional (ie, no covariates) model for a COA and an external measure (EM) in an LGM framework. In the example, i refers to an individual the COA was measured at four time points and the EM was measured at three time points, which do not have to be the same time points. Additionally, the model assumes a linear function of time. Figure 2 depicts this model using a path diagram.
In this example, the first dependent variable, the COA, was assessed at 4 time points (referred to as {Yi1, Yi2, Yi3, Yi4} at times {t1, t2, t3, t4} in Equations 1a–1d, also shown in the path diagram in Figure 2). These are regressed onto a latent intercept (BCOA0i), a latent slope (BCOA1
i), and the corresponding residual error {∊i1, ∊i2, ∊i3, ∊i4}. The second dependent variable, the EM, was assessed at 3 points (referred to as {Yi5, Yi6, Yi7} at times {t5, t6, t7} in Equations 1e–1g in Figure 2), and regressed onto a latent intercept (BEM0i), a latent slope (BEM1
i), and the corresponding residual error {∊i5, ∊i6, ∊i7}. The BCOA and BEM latent intercepts and slopes are assumed to be multivariate normally distributed, with mean structure and variance-covariance structure shown in Equation 2. The residual terms can be assumed independent and homoscedastic as in Equation 3 or may be parameterized to allow the residual terms of the BCOA and BEM to correlate. The longitudinal construct validity correlation between 2 dependent variables can be conceived as the correlation between their 2 latent slopes,

Path diagram representing a joint model of 2 dependent outcomes: clinical outcome assessment (COA) of interest and external measure used for validation purposes (EM). Note: Y1 to Y4 represent the observed COA data at 4 distinct time points; Y5 to Y7 represent the observed EM data at 3 distinct time points. The loadings of each Y on the corresponding intercept parameter (BCOA0 and BEM0) are fixed to 1; the loadings of each Y on the corresponding slope parameter (BCOA1 and BEM1) are fixed to the time coefficient (t1-t4 or t5-t7). For example, the time coefficients for Y measured at baseline, follow-up 1, follow-up 2, and follow-up 3 could be fixed at 0, 1, 2, and 3.
In exchange for the ability to statistically test the comparisons based on the a priori hypotheses, these techniques can require a sizeable number of observations (in order for the modeling algorithm to converge) and assume data that are missing at random. 20,24 Alternatively, when appropriate to ignore the time-varying nature of the relationships or to provide a signal of construct validity, the correlation values can be computed using change from baseline to the postbaseline treatment period average, for example, when the treatment response is immediate and consistent across time.
Known-Groups Analysis
Known-groups analyses are used to evaluate the discriminating ability of a COA by classifying study subjects into groups based on some characteristic with a hypothesized relationship to the construct and testing group mean differences. 1,10,41 –43 An optimal analysis will include hypotheses based on classifications that are strongly related to the COA’s construct. Evidence is weaker when group mean differences fail to achieve statistical significance (if hypothesized as strong) and weaker still when differences are in the wrong direction. Known-group differences for other measures provide an immediate context for the results computed for the target COA.
Reeve and colleagues’ survey of measurement professionals showed that only 41% would require known-groups analyses for support of construct validity while 57% felt it was “desirable but not required.” 5
Cross-sectional
Group-level score comparisons based on data from a single time point or a series of single time points can be conducted using t tests, analyses of variance (ANOVAs), or nonparametric tests depending on the nature of the COA scores and group classification variables.
Longitudinal
Group-level score comparisons can be conducted using repeated-measures ANOVA (when the known groups are constant) or the mixed-model framework (using the GLMM or the SEM framework when the known groups vary over time). 18,20,24,44 As with cross-sectional correlational analyses, the size and statistical significance support construct validity in the longitudinal case.
Ability to Detect Change
Necessary evidence for the longitudinal validity of a COA is its responsiveness—that is, its ability to detect change when change is expected (ie, the ability of a COA to measure a change in a clinical state that is valuable and meaningful to the patient or clinician). 1,2,43 The PRO Guidance favors information as to whether “there is evidence that PRO scores are affected by changes that are not specific to the concept of interest, the PRO instrument’s validity may be questioned.” 14 Evidence that the patient experience of a concept has changed in tandem with the COA of interest supports the COA’s responsiveness and its successful use in clinical trials. The PRO Guidance does not cite best methods for quantifying a COA’s ability to detect change. 14
An optimal hypothesis focuses on change over a period of time that includes a strong and relevant intervention (eg, an established therapy) or a duration that has been established as long enough for meaningful disease progression. 2 The nature of the medical condition must be taken into consideration when developing the responsiveness hypotheses. For example, in disorders such as oncology and muscle wasting, the treatment goals surround slowing or halting deterioration rather than facilitating improvements. Identifying a COA capable of detecting even the smallest deterioration is key in these types of conditions to understanding the disease and treatment impacts. Gathering evidence to support responsiveness can be more difficult in such circumstances compared with conditions in which quick and strong improvements can be achieved through established interventions. Computation of responsiveness statistics for other measures provides an immediate context for the size of the responsiveness statistic computed for the target COA.
The results of the ISOQOL survey by Reeve and colleagues 5 showed that a majority (57%) agreed that a measure to be used in longitudinal research should have evidence of responsiveness, and 42% considered this “desirable but not required.” A second question about responsiveness in the ISOQOL survey asked whether cross-sectional data were acceptable: “if a PRO measure has cross-sectional data that provide sufficient evidence in regard to the reliability, content validity, and construct validity but has no data yet on responsiveness over time, would you accept use of the PRO measure to provide valid data over time in a longitudinal study if no other PRO measure was available?” The use of cross-sectional data in lieu of longitudinal data was endorsed by 65% of the respondents, while 32% of respondents would require prior evidence of responsiveness. 5 Interestingly, in their review of sample sizes used to validate PRO measures, Anthoine and colleagues 45 reported that responsiveness was examined in only about 10% of the studies reviewed, typically via paired t test.
Cross-sectional
By definition, the evaluation of responsiveness requires at least 2 time points. In the absence of 2 time points, some researchers have suggested that a known-groups classification at a single time to compute a responsiveness statistic based on the assumption that the amount of cross-sectional difference between the 2 groups may in some circumstances be a suitable proxy for change over time. 46,47
Longitudinal
There is a large and extended family of responsiveness statistics, and a number of excellent sources describe and review measures of responsiveness. 4,46 –51 While Terwee and colleagues 51 identified more than 30 different statistics in the literature for quantifying a COA’s ability to detect change, many of these take the general form of a difference (between time points or groups) in the numerator and a measure of variability (eg, standard deviation [SD] or standard error [SE]) in the denominator, resulting in a difference or change value expressed, respectively, in standard deviation units or t test statistics, both of which are 2 types of standardized effect sizes.
From these various reviews, there are numerous ratios between a point estimate and a measure of variability for quantifying responsiveness. The more common and popular responsiveness statistics are the paired (or repeated measures) t test, effect size estimate of change ([Meanfollow-up – Meanbaseline] ÷ SDbaseline), standardized response mean ([Meanfollow-up – Meanbaseline] ÷ SDchange), standardized mean change difference ([Mean changegroup1 – Mean changegroup2] ÷ SDbaseline pooled), and Guyatt’s statistic 52 ([Meantreatment group at follow-up – Meantreatment group at baseline] ÷ SDchange in control). In addition to evaluating responsiveness using the family of statistics described in this section, the correlation between change in a COA and change in an external measure is a measure of responsiveness.
In a nutshell, these statistics provide a comparison of change based on a prespecified unit. The unit should be chosen based on the ultimate use of the COA. Deyo and colleagues 47 prefer a variant of Guyatt’s statistic, Norman and colleagues 50 recommend Cohen’s d and the standardized response mean (SRM), Zou 49 recommends the SRM, and Terwee and colleagues 51 recommend change correlations. Finally, the responsiveness statistic value for the target COA should exceed the observed statistic for another measure selected as a construct validity comparator, if available. 20
Cohen 30 provided multiple rules of thumb for interpreting effect sizes, including correlations (described earlier). Cohen’s rule of thumb regarding the interpretation of noncorrelation effect sizes is often-repeated and much-maligned: effect sizes of about 0.20 are small, those of about 0.50 are moderate or medium, and those ≥0.80 represent large effects. These simple guidelines are echoed by Kazis and colleagues, 53 Cappelleri and colleagues, 54 and others 55,56 : 0.2-0.49 are small, 0.5-0.79 moderate, and ≥0.80 are large. These are only rough guidelines and do not supplant a thorough familiarity with typical effect size results in the area of COA research and development. It is particularly important to compute effect size statistics for other measures included in the study to provide an immediate context for the size of the responsiveness statistic computed for the COA of interest. Note that disease-specific measures are generally expected to be more responsive than generic measures.
Discussion
Before using a COA to support decision making, research should evaluate the validity of the COA for the given population and intended purpose. If the planned assessment is related to change in the COA scores, it is imperative that the COA be shown to be a valid measure of change. Therefore, longitudinal evidence should be available. We have described a sample of methods for evaluating aspects of a COA’s longitudinal construct validity.
Beyond these, numerous additional methods offer some information about construct validity (eg, longitudinal item response theory). The selection of an assessment based on modeling is often highly dependent on the researcher’s training and experience. Biostatisticians may prefer to parameterize using mixed models, education researchers may prefer hierarchical linear models, and psychologists may focus on LGMs. As the numerical techniques and software packages evolve, even the subtle variations in terminology and parameterization across these models can make it difficult to bridge the differences and offer straightforward recommendations for how to apply these techniques to support longitudinal construct validity. For the interested researcher, we recommend Fitzmaurice and colleagues, 20,34 Davidian and Giltinan, 21 Grimm, 22 Stull, 23 Preacher, 24 Duncan and Duncan, 25 Bollen and Curran, 37 and Singer and Willet 26 as good resources to help build the technical background necessary to choose among approaches.
Standard practice should involve a thoughtful assessment with theory-driven hypotheses (ranked by importance or priority) using correlational analyses, known-groups analyses, and an evaluation of the measure’s ability to detect change. 19 The number and quality of the hypotheses (eg, availability of external measures that are highly relevant to the measured construct) and consistency of findings across methods can provide a substantial case for the COA’s longitudinal construct validity. 5 If a study includes few relevant measures, the strength of any inferences is diminished. Additional challenges for the collection of construct validity evidence include sample size, study design, and the level of missing data. 57 The PRO guidance emphasizes the COA’s fit for purpose, but psychometric evaluations using cross-sectional data pose a complication for small samples encountered in rare and orphan diseases. Longitudinal designs offer an opportunity to assess the psychometric properties of COAs for rare/orphan diseases. There is an important need for the COA literature to address small-sample longitudinal modeling, in particular, exploration of questionnaire structure, methods of scaling, and measures of goodness of fit.
In summary, although cross-sectional evaluations of construct validity provide some evidence in support of a COA intended to measure change (eg, a COA designed to show the benefit of an intervention), adding the complexity of the longitudinal data through multiple approaches may provide stronger evidence and should be considered when designing the evaluation. Selecting multiple external measures that assess both strongly similar and strongly dissimilar concepts will provide a stronger base of evidence, rather than using a limited set of external measures and potentially “double dipping.”
COAs allow us to measure patient-centered outcomes in a consistent way across studies, samples, and, where appropriate, across diseases. Because of the importance of identifying and developing appropriate items to measure these constructs, the development of a COA requires a large investment of time and resources. The subsequent pretesting and psychometric evaluation provide quantitative evidence to indicate whether the COA development was successful for its intended use. Evidence in support of construct validity—incorporating both cross-sectional data and longitudinal data—is a key component of the psychometric evaluation.
Footnotes
Acknowledgments
The authors acknowledge Sheri Fehnel, Donald Stull, and Lee Bennett for their review and technical contributions to this paper.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was performed and funded by RTI Health Solutions. All authors are employees of RTI Health Solutions.
