Abstract
The Hospital Anxiety and Depression Scale (HADS) measures anxiety and depressive symptoms and is widely used in clinical and nonclinical populations. However, there is some debate about the number of dimensions represented by the HADS. In a sample of 534 Dutch cardiac patients, this study examined (a) the dimensionality of the HADS using Mokken scale analysis and factor analysis and (b) the scale properties of the HADS. Mokken scale analysis and factor analysis suggested that three dimensions adequately capture the structure of the HADS. Of the three corresponding scales, two scales of five items each were found to be structurally sound and reliable. These scales covered the two key attributes of anxiety and (anhedonic) depression. The findings suggest that the HADS may be reduced to a 10-item questionnaire comprising two 5-item scales measuring anxiety and depressive symptoms.
The Hospital Anxiety and Depression Scale (HADS; see Table 1 for item content) is a brief, psychometrically sound instrument (Bjelland, Dahl, Haugh, & Neckelmann, 2002; Zigmond & Snaith, 1983) that has been used in diverse samples to assess symptoms of anxiety and depression (Bjelland et al., 2002; Herrman, 1997; Stafford, Berk, & Jackson, 2007). The HADS is largely robust across gender and age groups (Stordal et al., 2001). The scale consists of 14 items that contribute to two 7-item subscales measuring anxiety and depressive symptoms (score range 0-21; Zigmond & Snaith, 1983). The HADS is available in multiple languages (Bjelland et al., 2002), which together with the scale’s brevity make it feasible to include in multicenter, international studies.
Item Content, Mean Item Scores, Pattern Coefficients of the Three-Factor Model After Oblimin Rotation, and Results of the Automated Item Selection Procedure (AISP) in Exploratory Mokken Scale Analysis
Note. us = unscalable. Boldface values indicate per item the highest pattern coefficient for the three-factor model.
The sensitivity and specificity of the HADS as a screening tool for probable anxiety and depressive disorder are acceptable (Bjelland et al., 2002; Stafford et al., 2007). Its overall screening properties for depressive disorder are comparable with that of the Patient Health Questionnaire (PHQ-9; Stafford et al., 2007). The PHQ-9 was developed to mirror the Diagnostic and Statistical Manual of Mental Disorders, fourth edition (DSM-IV; American Psychiatric Association, 1994) criteria for a clinical diagnosis of depression on an item-to-item basis (Kroenke, Spitzer, & Willians, 2001). An advantage of using the HADS in the context of somatic disease is that the scale does not contain items that reflect somatic indicators of depression and anxiety (e.g., sleeplessness, racing heart, fatigue), which may be the result of patients’ medical condition rather than psychological distress. Inclusion of somatic indicators could falsely increase the prevalence rate of anxiety and depression in such patient groups (Herrman, 1997; Stafford et al., 2007), which should obviously be avoided.
The HADS is a commonly used outcome measure in cardiac populations and has been employed in studies of patients treated with percutaneous coronary intervention (PCI; Pedersen, Denollet, Van Gestel, Serruys, & Van Domburg, 2008), implantable cardioverter defibrillator therapy (Spindler, Johansen, Andersen, Mortensen, & Pedersen, 2009; Undavia et al., 2008), peripheral arterial disease (Smolderen et al., 2009), and chronic heart failure (Schiffer et al., 2005). The HADS has also been used as a determinant, showing that anxiety enhances the negative effect of depression on health status in PCI patients (Pedersen et al., 2006), and depressive symptoms predict recurrence of atrial fibrillation after cardioversion (Lange & Herrmann-Lingen, 2007). Preliminary evidence from studies on patients with coronary artery disease (Doyle, McGee, De La Harpe, Shelley, & Conroy, 2006; Herrman, Brand-Driehorst, Buss, & Rüger, 2000), chronic heart failure (Jünger et al., 2005), and cancer (Groenvold, Petersen, Idler, Bjorner, & Mouridsen, 2007; Grulke, Larbig, Kächele, & Bailer, 2008) also shows that psychological distress, as measured with the HADS, predicts survival.
There is some discussion in the literature, however, whether the HADS measures two distinct facets of mood, as originally proposed by Zigmond and Snaith (1983). In a recent review of the HADS, the two-factor structure was confirmed in the majority of studies (Bjelland et al., 2002; see also Mykletun, Stordal, & Dahl, 2001; Spinhoven et al., 1997). However, some authors use a total HADS score (Scherer et al., 2006), despite the model fit for a single factor being poor (Dunbar, Ford, Hunt, & Der, 2000), whereas others advocate a three-factor structure. Dunbar et al. (2000) applied confirmatory factor analysis (CFA) to 2,547 participants from three age cohorts. Their findings support the “tripartite” model of anxiety and depression, as proposed by Clark and Watson (1991), comprised of negative affectivity, anhedonic depression, and autonomic anxiety. The notion is that anxiety and depression have distinctive features, with low positive affect/anhedonia (i.e., loss of pleasure and interest in life) being distinct to depression, and autonomic arousal marked by somatic symptoms (i.e., feelings of panic and trembling) being distinct from anxiety, and with negative affectivity underpinning both facets. Other investigators found further support for a three-factor model (Caci et al., 2003; Friedman, Samuelian, Lancrenon, Even, & Chiarelli, 2001; Martin, Lewin, & Thompson, 2003; Martin, Thompson, & Barth, 2008; Rodgers, Martin, Morse, Kendell, & Verril, 2005). The three-factor models reported in these studies, however, differ in their factorial composition (Caci et al., 2003; Dunbar et al., 2000), but none of the studies revealed compelling evidence for preferring one particular three-factor model to others. For the use of the HADS in research and clinical practice, knowledge of its exact dimensionality is important because the dimensionality determines whether one, two, or three total scores are needed for the proper interpretation of a patient’s psychological state.
When trying to settle the issue of the number of factors of the HADS, most studies have used principal components analysis (PCA) and CFA (e.g., Bjelland et al., 2002). Recently, Mokken scale analysis (MSA; Mokken, 1971; Sijtsma & Molenaar, 2002) has become popular in the medical context for assessing the dimensionality of questionnaire data in addition to PCA and factor analysis methods. Up-to-date examples of medical MSA applications include Moorer, Suurmeijer, Foets, and Molenaar (2001), Meijer and Baneke (2004), Michielsen, De Vries, Van Heck, Van der Vijver, and Sijtsma (2004), Roorda et al. (2005), Ivarsson and Malm (2007), Valenzuela and Sachdev (2007), Korner et al. (2008), Sijtsma, Emons, Bouwmeester, Nyclíček, and Roorda (2008), Stochl, Boomsma, Van Duijn, Brozova, and Ruzicka (2008), Watson, Deary, and Shipley (2008), and Bech, Wilson, Wessel, Lunde, and Fava (2009).
Compared with PCA and CFA, MSA has several advantages (e.g., see Wismeijer, Sijtsma, Van Assen, & Vingerhoets, 2008). First, like PCA and CFA, MSA assesses dimensionality, but unlike PCA and CFA, it also uses modern methods from nonparametric item response theory (IRT; e.g., Sijtsma & Molenaar, 2002) to test the psychometric properties of unidimensional scales that are found. Second, the assumptions of the underlying IRT model are assessed, which results in either the support or the rejection of the psychometric properties for the scale. PCA does little more than compute the eigenvalues of the interitem correlation matrix but has no underlying measurement model. Like PCA, CFA does not rest on a measurement model, but like IRT, CFA explicitly tests an underlying dimensionality model. Third, MSA is particularly suited for analyzing discrete questionnaire data, for example, stemming from Likert-type items such as those used in the HADS, whereas PCA and CFA are suited for continuous data. Thus, MSA avoids “difficulty factors” and distortions due to skewed item-score distributions; PCA and CFA often use tetrachoric or polychoric correlations, which assume latent normal distributions for highly discrete item scores. This is a restrictive assumption, which may also lead to distorted results (Hattie, 1985).
MSA also has several advantages over parametric IRT models, such as the rating scale model (e.g., Andrich, 1978). First, MSA is based on less restrictive assumptions about the data than parametric item response models (e. g., Embretson & Reise, 2000) while maintaining important measurement properties. This prevents researchers from unnecessarily removing items from a scale that have good measurement properties but do not satisfy the constraints imposed by the more restrictive parametric models. Second, MSA provides useful tools for exploratory dimensionality analysis that are not readily available for parametric IRT models, rendering MSA particularly suited for dimensionality analysis of the HADS.
The objectives of the current study were (a) to briefly introduce MSA as a promising and fruitful method for the analysis of questionnaire data collected in psychological research and clinical assessments, (b) to examine the dimensionality of the HADS using MSA and contrast findings with factor analysis, and (c) to assess the scale properties of the HADS in a sample of cardiac patients treated with PCI.
Mokken Scale Analysis
MSA is a psychometric methodology, and the purpose here is to briefly discuss it in a manner that is comprehensible to nonstatisticians. MSA refers to a set of psychometric data analysis methods that assess the fit of two IRT models to questionnaire data. These IRT models are the monotone homogeneity model (MHM) and the double monotonicity model (DMM); see Sijtsma and Molenaar (2002) and Sijtsma and Meijer (2007) for extensive discussions of the models. The MHM is important because it justifies the ordering of individuals on a latent variable by means of their total scores based on the items in the questionnaire. For dichotomous items (i.e., 0/1 scoring), the DMM is important because it implies an ordering of items, which is invariant for all scale values. In the clinical–medical assessment literature and elsewhere, such an invariant item ordering (Sijtsma & Hemker, 1998; Sijtsma & Junker, 1996) is sometimes called a hierarchical scale or a cumulative scale. However, for polytomous items (i.e., three or more ordered item scores; the HADS items have four ordered scores), Sijtsma and Hemker (1998) proved that the DMM does not imply a hierarchical scale. Consequently, whether a set of polytomous items forms a hierarchical scale should be investigated independently of the DMM (Van der Ark, 2001), and for this purpose Ligtvoet, Van der Ark, Te Marvelde, and Sijtsma (2010) proposed an effective methodology.
In this study, we use both the MHM and the Ligtvoet et al. (2010) methodology for investigating whether a set of polytomous items constitute a hierarchical scale to evaluate the psychometric properties of the HADS. In this section, first we discuss the MHM and then the hierarchical-scale methodology.
Mokken Scale Analysis for the Monotone Homogeneity Model
Assumptions of the MHM
The MHM assumes that a set of items meant to be included in the same scale measure the same attribute, such as depression or anxiety, and that each item has a positive and monotone relation with this attribute. The first assumption, known as unidimensionality, is obvious and prevents the total score based on the items to represent a hodgepodge of different attributes. Suppose one would allow different items in one scale to measure anxiety, neuroticism, depression, introversion, extraversion, and dominance. Consequently, a patient’s total score would be incomprehensible because it would represent a complex mixture of attributes, and the total scores of different patients would be incomparable because their composition would likely be different depending on their specific individual item scores. For example, for two patients with the same total score, one score could primarily reflect high scores on neuroticism and extraversion items, and another identical score could reflect high item scores on depression and introversion items. Comparing these scores would be meaningless.
The second assumption, known as monotonicity, stipulates that a higher attribute level corresponds to a higher expected item score. For an HADS depression item, a higher depression level is expected to correspond to a higher score on the item. As an illustration, Figure 1A shows three monotone functions, known as item response functions (IRFs), that display the relationships of the cumulative scores 1, 2, and 3 on an item having four answer categories (scored 0-3) with the attribute of interest. The highest function represents the probability that someone with a particular attribute score (horizontal axis) has at least a score of 1 on the item; that is, either a score 1, 2, or 3. The middle function represents the probability of at least an item score of 2, and the lowest function a probability of exactly 3. Every person has a score of at least 0, so a function representing the corresponding probability equals 1 irrespective of the attribute value and is not informative for assessing the item’s functioning.

(A) Example of item with three IRFs (four item scores) relating response probability (x ≥ 1, 2, 3) to attribute scale. (B) Two sets of IRFs for two items agreeing with the DMM, which have intersecting summary IRFs. (C) Estimated IRFs (solid) and summary IRF (dashed-dotted; adapted from MSP). (D) Estimated continuous summary IRF (solid) and 90% confidence interval (dashed; adapted from TestGraf) for item A4 from the anxiety scale.
Monotonicity is an intuitively appealing assumption but also a technically essential one, because together with the unidimensionality assumption (and a technical assumption known as local independence, which we ignore here to keep the discussion nontechnical) it guarantees that persons can be ordered on the scale for the attribute by means of their total scores when items are dichotomous (Hemker, Sijtsma, Molenaar, & Junker, 1997). The assumptions strongly support this ordering when items are polytomous (Van der Ark, 2005), although small but practically unimportant deviations are possible. It is important to note that such an ordinal scale for an attribute is not simply obtained by summing an individual’s scores on the questionnaire’s items (this is just counting); on the contrary, it has to be shown by means of psychometric analysis (a) that the items measure the same attribute and (b) that the items are monotonically related to this attribute. This is exactly what scaling method MSA does.
Constructing Scales
In MSA for the MHM, for item j the item scalability coefficient Hj summarizes the strength of the relationship between an item and the attribute scale. Given unidimensionality and monotonicity, it can be shown (e.g., Sijtsma & Molenaar, 2002, p. 59) that 0 ≤ Hj ≤ 1. The higher the Hj value, the better the item separates low attribute total scores from high attribute total scores (i.e., the item has high discrimination power). In practical data analysis, a generally accepted rule of thumb for items to be included in a scale is Hj ≥ .3, so that only those items are included that have at least moderate discrimination power (Sijtsma et al., 2008).
For the set of items comprising one scale, the scalability coefficient H uses the information from the item scalability coefficients Hj to summarize the strength of the relationship between the total score and the attribute scale. A higher coefficient H reflects a more accurate person ordering on the attribute scale by means of the total score (Sijtsma & Molenaar, 2002, p. 60). Given unidimensionality and monotonicity, it can be shown that 0 ≤ H ≤ 1, and for scales to be used in practice, the following rules of thumb are used: .3 ≤ H < .4 means a weak scale; .4 ≤ H < .5 a medium scale; H ≥ .5 a strong scale; and H < .3 means that the items are unscalable (Mokken, 1971, p. 185). These rules of thumb ascertain whether a set of items imply an accurate ordering of persons on an attribute scale defined by the items (also, see Mokken, Lewis, & Sijtsma, 1986). An MSA including scalability evaluation using H, however, does not assess whether the items constitute a hierarchical scale. This topic is discussed shortly.
When a questionnaire contains items all measuring the same attribute, MSA can be used in a confirmatory way, but what if a questionnaire consists of items measuring different attributes, and the true dimensionality is unknown or liable to dispute? This is the case for the HADS. MSA provides an automated item selection procedure (AISP; for details, see Mokken, 1971, pp. 190-194; Sijtsma & Molenaar, 2002, chap. 5) for finding different clusters of items in an exploratory way, each cluster measuring a different attribute. The programs MSP (Mokken scale analysis for polytomous items; Molenaar & Sijtsma, 2000) and mokken1.4 (Van der Ark, 2007) include the AISP, which works as follows. The AISP selects item clusters, such that particular selection criteria are satisfied. This is done in an item-by-item step procedure, starting with the two items that have the highest significant H value that exceeds c (c ≥ 0), and then adding items one by one on the basis of the magnitude of their Hj value with respect to the items already selected into the cluster. When one cluster is selected and items are left unselected, the AISP tries to select from the remaining items a second cluster, a third cluster, and so on, until no items are left or only items are left that do not satisfy the criteria for inclusion. The researcher can manipulate the strictness of the selection by specifying an acceptable c value for scalability coefficients; the higher c, the stricter the procedure and the smaller the selected item clusters (provided, there are any). Based on the rules of thumb for coefficient H, MSP and mokken1.4 use c = .3 as the default lower bound but the researcher may use a different value when deemed necessary.
Hemker, Sijtsma, and Molenaar (1995; see also Sijtsma & Molenaar, 2002, pp. 80-86) recommend investigating dimensionality by running the AISP consecutively using c values such that c = .00, .05, . . ., .55. Different patterns of outcomes suggest different dimensionality solutions (Sijtsma & Molenaar, 2002, p. 81). As c increases, unidimensionality is apparent from (a) most or all items in one scale; (b) items leave the scale, which results in one smaller scale; and (c) the scale becomes smaller or splits into a few small scales, while many items are excluded. The typical pattern of results for multidimensionality is, as c increases, (a) most or all items are in one scale; (b) two or more scales are formed; and (c) items leave the scales formed in Step 2 and become smaller, or the scales fall apart into many subscales with only a few items per subscale, while many items are excluded. The crucial difference is in the second phase, when either one or multiple substantial scales are formed.
Data Analysis Steps in MSA for the MHM
A typical MSA for the MHM involves three data analysis steps, which can be performed using the program MSP (Molenaar & Sijtsma, 2000; for information on MSP, contact the first author) and the R program mokken1.4 (Van der Ark, 2007); R is for free and downloadable from http://cran-mirror.cs.uu.nl/). The steps are the following.
Step 1: Investigating dimensionality
Two possibilities exist. Exploratory dimensionality analysis involves finding unidimensional clusters of items from a larger pool of items relying on statistical outcomes and without committing oneself a priori to a particular dimensional structure. Confirmatory dimensionality analysis evaluates whether one or more a priori identified sets of items, which serve as null hypotheses, can be considered unidimensional.
Step 2: Investigating monotonicity
Selected items may have IRFs showing local decreases not revealed by the Hj values even when these values exceed the lower bound c. Such violations of monotonicity may distort the rank ordering of persons on the attribute scale by means of the total score. MSP and mokken1.4 test observed decreases in IRFs for significance. Another program, called TestGraf98 (Ramsay, 2000; downloadable for free from http://www.psych.mcgill.ca/faculty/ramsay/TestGraf.html) estimates summary IRFs, which wrap up the information from the different IRFs for one item. Also, see Ramsay (1991) and Junker and Sijtsma (2000).
Step 3: Scale properties
For unidimensional scales with monotone IRFs, the reliability of the total score is estimated using a method proposed by Sijtsma and Molenaar (1987), designed especially for MSA applications, and usually closer to the true reliability than lower bound Cronbach’s alpha (method included in MSP). Another possibility is to estimate the information function, which shows measurement precision depending on the scale value (included in TestGraf98). This is expected to be useful for the skew total score distributions of the HADS. Both reliability and the information function are important for evaluating the usefulness of the HADS for screening of individual patients for depression disorders.
Mokken Scale Analysis for Identifying Hierarchical Scales
Like the MHM, the DMM assumes unidimensionality (and local independence, which was ignored here) and monotonicity. Figure 1B shows the typical additional assumption of the DMM, which is that IRFs of different items do not intersect. For dichotomous items, scored 0 (e.g., disagree) and 1 (e.g., agree), there would only be one solid line and one dashed line, both representing the probability of having a 1 score, and which do not intersect under the DMM. Because these probabilities have the same ordering for all scale values, the items are invariantly ordered, thus providing a hierarchical scale. However, for polytomous items the situation is totally different because of the following aggregation problem.
An invariant item ordering refers to the items as entities and not to individual IRFs as shown in Figure 1B, which are informative only about particular item scores, as expressed by the probabilities of having a score in excess of x. However, what one wants to know in practical psychometric analysis is whether items are invariantly ordered, and not whether the individual IRFs of different items are invariantly ordered. The invariant item ordering can be investigated as follows (Ligtvoet et al., 2010). For each item, the individual IRFs are replaced by their mean function, which is the mean item-score conditional on the scale values. Figure 1B shows these summary IRFs as dotted curves. Whether a set of items has an invariant ordering can be investigated by checking whether the summary IRFs intersect or not. In Figure 1B, the summary IRFs intersect at attribute value −0.025; to the left of this point the “solid” item is less popular than the “dashed” item, and to the right the ordering is opposite. Hence, the items are not invariantly ordered.
Based on Figure 1B, we conclude that a fitting DMM is insufficient for an invariant item ordering or, equivalently, a hierarchical scale, and that the summary IRFs have to be investigated for drawing conclusions about item ordering; see Ligtvoet et al. (2010) for technical details. The computations can be done using program mokken1.4. Readers who are interested to learn more about the surprising result that for polytomous items the DMM does not imply an invariant item ordering may consult Sijtsma and Hemker (1998), who provide the mathematical proof. The key to understanding these results lies in realizing that almost all IRT models leave the mutual ordering of the IRFs of different items free, thus imposing too few restrictions to imply an invariant item ordering based on the summary IRFs. Sijtsma and Hemker (1998) also prove that, like the DMM, popular IRT models for polytomous items, such as the partial credit model (Masters, 1982) and the graded response model (Samejima, 1969) do not imply an invariant item ordering. The rating scale model (Andrich, 1978) provides an exception, but this model is more restrictive than the methodology that Ligtvoet et al. (2010) propose, and hence may lead to the unnecessary removal of items even when they fit well into an invariant ordering with other items.
We propose an MSA, which includes assessing the fit of the MHM and investigates whether the items in the resulting scales have an invariant ordering, hence constituting a hierarchical scale. The four steps for identifying hierarchical Mokken scales are (a) investigating dimensionality, (b) investigating monotonicity, (c) investigating invariant item ordering, and (d) assessing the scale properties. Thus, the difference with the MHM analysis resides in the addition of the third step.
Mokken Scale Analysis and Factor Analysis of HADS Data
In what follows, we use the four steps of MSA to analyze the HADS. Because factor analysis is the most frequently used method for dimensionality analysis, we compare our MSA results to results from both exploratory factor analysis (EFA) and CFA.
Method
Participants
A consecutive cohort of cardiac patients treated with PCI using the paclitaxel-eluting stent as the default strategy, recruited from the Erasmus Medical Center, Rotterdam, the Netherlands, between July 1, 2003, and July 1, 2004, completed the HADS at baseline (i.e., 4 weeks after PCI). The medical ethics committee of the hospital approved the study. All patients provided written informed consent, and the study was conducted in agreement with the Helsinki Declaration. Details of the study design have been previously published (Pedersen et al., 2007).
Out of 845 patients, 19 patients died within the first 4 weeks after the index procedure, and 116 patients were excluded because of insufficient knowledge of the Dutch language. The remaining surviving patients (n = 710) were approached in writing and asked to complete the HADS 4 weeks post-PCI, with 536 patients (75%) agreeing. Compared with responders, excluded patients and nonparticipants (n = 309) were more likely to smoke (22% versus 14%; χ2(1) = 8.95, p =.003) but were less likely to suffer from dyslipidemia (63% versus 74%; χ2(1) = 11.23, p = .001). No other differences were found between excluded patients/nonresponders and responders on baseline characteristics, including cardiac medication.
Measures
The HADS is a 14-item self-report measure, with 7 items contributing to an anxiety subscale (e.g., “I feel tense or wound up”) and 7 items to a depressive symptom subscale (e.g., “I have lost interest in my appearance”; Zigmond & Snaith, 1983). Items are answered on a 4-point Likert-type scale running from 0 to 3, and total score ranges run from 0 to 21 for both subscales (Zigmond & Snaith, 1983). We used the Dutch version of the HADS, which was validated by (Spinhoven et al., 1997) in a mixed population of 6,165 individuals, including younger adults, elderly persons aged 57 to 65 years, elderly persons 66 years or older, patients seen in general practice, outpatients with unexplained somatic symptoms, and psychiatric outpatients.
Procedure
Two patients did not fill in the HADS, leaving a sample size of 534 patients (72.5% men; mean age = 63.29, SD = 11.03 years). Ten respondents had one or two missing item scores on the HADS-A (anxiety) scale, and 135 respondents had one or two missing item scores on the HADS-D (depression) scale. For these respondents, missing values were imputed by means of two-way imputation (Van Ginkel & Van der Ark, 2005; Van Ginkel, Van der Ark, & Sijtsma, 2007). Imputation did not have a significant effect on the distributional characteristics of the total scores. A sample size of 534 respondents is considered adequate for MSA and factor analysis (Pett, Lackey, & Sullivan, 2003).
Data analysis was done as follows:
We used SPSS 16.0 (SPSS Inc., 2008) for EFA on the HADS data, and discussed the factorial composition found. We used principal axis factoring and the scree plot (Cattel, 1966; Pett et al., 2003, pp. 118-120) and Kaiser’s eigenvalue-greater-than-1 criterion for factor extraction, and Oblimin rotation to obtain interpretable factor solutions.
We used AMOS 7.0 (Arbuckle, 2006) for CFA to test five competing factor models for the HADS, as discussed in the HADS literature. We used the comparative fit index (CFI; Bentler, 1980) and the root mean square error of approximation (RMSEA; Browne & Cudeck, 1993) to assess goodness-of-fit of factor models. Kline (2005, pp. 139-140) proposed the following rules of thumb: CFI < .90 indicates poor fit, .90 < CFI < .95 indicates reasonable fit, and CFI > .95 indicates good fit; and RMSEA > .10 indicates poor fit, .05 < RMSEA < .10 indicates reasonable fit, and RMSEA < .05 indicates good fit.
We combined the results of EFA and CFA and discussed the scale properties for the best factorial solution. Among these properties were three estimates of the total-score reliability (Sijtsma, 2009), which were Cronbach’s alpha, Guttman’s lambda2 (both computed using SPSS 16.0), and the greatest lower bound (GLB; computed using program MRFA2, Ten Berge & Kiers, 2003; program downloadable from http://www.ppsw.rug.nl/~kiers/).
3. We used MSP (Molenaar & Sijtsma, 2000) for exploratory MSA, following the first two steps of the four-step methodology, thus investigating dimensionality and monotonicity. We used the program mokken1.4 (Van der Ark, 2007) for assessing invariant item ordering.
4. We used MSP and mokken1.4 to do confirmatory MSA for the same five competing factorial models for the HADS that were also tested using CFA, using the first three steps of the four-step methodology.
Finally, we combined the results of exploratory and confirmatory MSA, and discussed the scale properties for the dimensionally best solution. We used MSP to estimate total-score reliability by means of the method proposed by Sijtsma and Molenaar (1987), and TestGraf98 (Ramsay, 2000) to estimate local measurement precision for total scores.
Results
Exploratory Factor Analysis
Both inspection of the scree plot (Figure 2) and Kaiser’s eigenvalue-greater-than-1 criterion led to the extraction of three factors, accounting for 36.5% (Factor 1), 7.1% (Factor 2), and 3.5% (Factor 3) of the common variance. Oblimin rotation was used to produce a three-factor solution of which the pattern coefficients (Table 1, columns 4-6) provided indications of overlapping depression, anxiety, and negative affect attributes (Clark & Watson, 1991). Five items measuring anxiety loaded highest on the first factor, six items measuring depression loaded highest on the second factor, and two items measuring anxiety and one item measuring depression loaded highest on the third factor.

Scree plot from exploratory factor analysis of the Hospital Anxiety and Depression Scale data
Confirmatory Factor Analysis
CFA was used to test five different factor models (see Table 2, columns 1-4). Ravazi, Delvaux, Farvacques, and Robaye (1990) and Smith et al. (2006) proposed a one-factor model for the HADS, which assumes that the 14 items measure the same factor. This model complies with the practice of using a total score based on all 14 items. Alternatively, the original authors of the HADS (Zigmond & Snaith, 1983) proposed a two-factor model, representing different factors for anxiety and depression. Dunbar et al. (2000) proposed a three-factor model in which item A4 (“I can sit at ease and feel relaxed”) loads on two factors. They also assumed an additional unique covariance between items A6 (“I feel restless”) and D7 (“I can enjoy a good book”), which cannot be explained by the correlation of the underlying factors on which the items are assumed to load. This means that the item clusters in Dunbar et al.’s model are not unidimensional. Friedman et al. (2001) proposed a three-factor model based on unidimensional clusters of items, whereas Caci et al. (2003) proposed a different three-factor model. The item clusters are given in Table 2 (column 4). The one-factor, two-factor, and 3 three-factor models were all tested in the current study.
Factor Structure and Fit Indices (CFI and RMSEA) of a CFA for the One-Factor, Two-Factor, and the 3 Three-Factor Models
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; CFA = confirmatory factor analysis.
In the Dunbar et al. (2000) model, a correlation between the residuals of Items A6 and D7 is assumed.
Caci et al. (2003) proposed a second model in which Item D5 was removed from the Depression scale. The CFI and RMSEA for this model were .946 and .065, respectively.
Based on the CFI and the RMSEA, the one-factor model fitted poorly (Table 2, columns 5-6), and was rejected. The two-factor model fitted by approximation. For the 3 three-factor models, the Caci et al. (2003) model showed the best fit (highest CFI and smallest RMSEA), but the differences in fit with the other 2 three-factor models were small.
Scale Analysis Results for EFA and CFA
EFA (Table 1) and CFA (Table 2) led to the conclusion that the Caci et al. three-factor model fits best. CFA analysis of the Caci et al. (2003) three-factor model showed that for anxiety the standardized regression weights (not tabulated) ranged from .63 (Item A5) to .76 (Item A2). For depression, the lowest standardized regression weight was .49 (Item D5), which was relatively small compared with the weights of the other items; these weights ranged from .59 (Item D4) to .79 (Item D2). For restlessness, the standardized regression weights ranged from .51 (Item A6) to .73 (Item A4). The correlation between the depression and restlessness factors was .62; between restlessness and anxiety .68; and between depression and anxiety .66. Table 3 shows reliability results for the three total scores. As expected, lambda2 improved little on alpha but GLB showed greater improvement.
Scale-Score Reliability for Caci et al.’s (2003) Three-Factor Model, using Cronbach’s Alpha, Guttman’s Lambda2, the Greatest Lower Bound (GLB), and the Mokken Method
Exploratory Mokken Scale Analysis
Step 1: Investigating dimensionality
Table 1 (columns 7-12) gives the results of the AISP for c = .0, .3, .4, and .5 (other c values are excluded here to prevent tedious repetition of almost similar results). For c = .0, all items were clustered into one scale. For c = .3, item A6 (“I feel restless as if I have to be on the move”) proved unscalable (in the table, denoted as us), whereas the other items were clustered into one scale. For c = .4, two scales were found: The first scale consisted of 10 items (Table 1, column 9), the second scale of the two items A4 (“I can sit at ease and feel relaxed”) and D7 (“I can enjoy a good book or radio or TV program”), whereas the two remaining items were unscalable. For c = .5, the 10-item scale found for c = 0.4 separated into two five-item scales, whereas the other four items were unscalable. Inspection of the cluster pattern obtained across the different c values (Sijtsma & Molenaar, 2002, p. 81) suggests that the HADS contains either one medium 10-item scale or two strong 5-item scales.
Step 2: Investigating monotonicity
Based on the dimensionality results, we investigated monotonicity for the 10-item scale (Table 1; c = .4) and the two 5-item scales (Table 1; c = .5). For each item, we counted the number of significant decreases in its IRFs, and determined an item summary value denoted by Crit (Molenaar & Sijtsma, 2000, p. 74): Crit < 40 means that violations of monotonicity may be ascribed to sampling error; 40 ≤ Crit < 80 indicates mild violations of monotonicity; and Crit ≥ 80 indicates serious violations of monotonicity. For the 10-item scale and the two 5-item scales, none of the items had Crit values exceeding 40. Thus, in each scale the monotonicity assumption held for each of the items.
Step 3: Investigating invariant item ordering
Table 4 (upper panel) shows the results for invariant item ordering for the 10-item scale and the two 5-item scales. For the 10-item scale, 8 items showed sample violations of invariant item ordering, and for 6 items a few violations were significant (α = .05; one-tailed test). Thus, the 10 items do not constitute a hierarchical scale. For the two 5-item scales, one scale (Cluster 1) had 3 items showing significant violations of invariant item ordering; hence, invariant item ordering was rejected. The other scale (Cluster 2) did not have significant violations. The corresponding HT value was .35, which indicates that the invariant item ordering had low accuracy (Ligtvoet et al., 2010). Thus, firm conclusions about the hierarchy of the items are not justified.
Invariant Item Ordering Results for Exploratory MSA (Upper Panel) and Confirmatory MSA (Lower Panel) for the One-Factor, Two-Factor, and the 3 Three-Factor Models
Confirmatory Mokken Scale Analysis
Step 1: Investigating dimensionality
Table 5 (column 2) shows the results for the unidimensional model. Under MSA, all item Hj values were positive, but Item A6 failed to exceed the practical lower bound criterion of c = .3. Including this item, the total-scale H value equaled .39, which indicates a weak scale. For the two-dimensional model (Table 5, columns 3 and 4), all item Hj values exceeded c = .3. The total-scale H values equaled .46 for anxiety and .45 for depression, in both cases indicating medium scales. For the Dunbar et al. (2000) three-factor model (Table 5, columns 5-7), the total-scale H values were .58 for autonomic anxiety, .43 for depression, and .43 for negative affectivity. For the Friedman et al. (2001) three-factor model (Table 5, columns 8-10), the total-scale H values were .60 for psychic anxiety, .51 for depression, and .42 for psychomotor agitation. Finally, for the Caci et al. (2003) three-factor model (Table 5, columns 11-13), the total-scale H values were .59 for anxiety, .51 for depression, and .40 for restlessness. For anxiety, all item Hj values exceeded .53. For depression, Item D5 (“I have lost interest in my appearance”) had a substantially lower item-scalability value than the other items in the scale. Excluding Item D5 from the depression scale as suggested by Caci et al. (2003) resulted in a total-scale H value of .57. Confirmative MSA suggested both Friedman et al.’s (2001) and Caci et al.’s (2003) models as possible candidates for describing the dimensionality structure.
Confirmatory MSA (H Values) for the One-Factor, Two-Factor, and the 3 Three-Factor Models
Note. Anx = anxiety; Depr = depression; Auton Anx = autonomic anxiety; Anhed Depr = anhedonic depression; Neg Affect = negative affectivity; Psych Anx = psychic anxiety; PsyMot Agit = psychomotor agitation; Restl = restlessness.
Caci et al. (2003) proposed a second model in which Item D5 was removed from the Depression scale. The H value for this scale was .57.
Step 2: Investigating monotonicity
The monotonicity assumption was investigated for the scales of the one-, two-, and the three-factor models. For the one-factor model, several items showed significant violations of monotonicity, but only the items A6 and D7 had Crit values in excess of 80, indicating misfit. Thus, the usefulness of the total score on all 14 items for rank ordering individuals may be questioned.
For the two-factor model, inspection of the results for the anxiety scale showed Crit = 44 for Item A4 (“I can sit at ease and feel relaxed”), which is a borderline case and thus difficult to interpret; hence, monotonicity of the three IRFs and the summary IRF of Item A4 were graphically examined (Figures 1C and 1D). MSP was used to estimate the three IRFs (Figure 1C, solid curves) and the summary IRF (dashed-dotted curve; represents mean conditional item score, rescaled between 0 and 1). The largest decrease (equal to .18) was found between rest-score groups 4 and 5 of the estimated IRF for an item score of at least 1. The summary IRF also showed a small decrease, suggesting that monotonicity was also violated mildly at the item score level. TestGraf98 was used to estimate the summary IRF by means of kernel smoothing (Figure 1D), but the graph did not corroborate the local decrease found by means of MSP. We conclude that monotonicity is tenable for Item A4.
For the other items in the anxiety subscale, Crit values were smaller than 40. Also, for the depression items under the two-factor model, all Crit values were smaller than 40. Thus, the conclusion is that for all anxiety and depression items monotonicity held, allowing a person ordering using total score X+.
For the Dunbar et al. (2000) three-factor model, depression item D5 showed four significant violations, and its Crit value was 97. None of the items from the other two scales showed significant violations of monotonicity. For the other 2 three-factor models, all Crit values were below 40. Thus, scales derived under these 3 three-factor models have monotone IRFs.
Step 3: Investigating invariant item ordering
We investigated invariant item ordering for all five factorial models (Table 4; lower panel). Each of the scales of the Ravazi et al. (1990) one-factor model and the Zigmond and Snaith (1983) two-factor model had significant violations, thus rejecting an item hierarchy. For the Dunbar et al. (2000) three-factor model, only the scale negative affectivity constituted a hierarchy but the HT value suggested low accuracy. For the Friedman et al. (2001) three-factor model, a hierarchical scale of low accuracy was found for psychomotor agitation. For the Caci et al. (2003) three-factor model, the anxiety and depression scales each had violations; hence, these scales did not have an invariant item ordering. The restlessness items constituted a hierarchical scale but the HT value indicated that the item ordering was inaccurate. To summarize, some scales were hierarchical but the statistical evidence for this conclusion was weak.
Step 4: Scale properties: Results for exploratory and confirmatory MSA
Exploratory and confirmative MSA led to the conclusion that the Caci et al. (2003) model may be preferred. Two of the three scales found in exploratory MSA, anxiety and restlessness, had the same composition as the factors in the Caci et al. model. Item D5 from the depression scale in the Caci et al. model was found to be unscalable in the exploratory MSA when one would only accept strong scales in the final outcome, but scalable when also weak and medium scales would be acceptable. The confirmatory MSA of the depression scale yielded HD5 = .39, which is well over the practical lower bound c = .3, and there were no serious violations of the monotonicity assumption. Based on these results, Item D5 may be included in the depression scale. Thus, the final scale measuring depression consists of six items.
All anxiety items were strong indicators (Hj ≥ .53 for all items; see Table 5, column 11). The reliability of the anxiety scale was .84 (Table 3, last column). Confirmatory MSA showed that depression item D5 had a somewhat lower Hj value (HD5 = .39) than the other depression items (Table 5, column 12). This item is a weak indicator of the underlying attribute, also contributing weakly to the rank ordering of persons on depression. The other item Hj values ranged from .48 to .60. The reliability of the depression scale was .84 (Table 4, last column). For the restlessness scale, the strongest indicator was Item A4 (“I can sit at ease and feel relaxed”; Table 5, column 13). The reliability was .63 (Table 3, last column). The reliability of the scale is high enough to compare group means, but too small for individual decision making.
Figure 3 shows local measurement precision (solid line) as a function of total score for the Caci et al.’s (2003) three-factor model (graphs adapted from TestGraf98), and the total-score distribution for each scale. For the anxiety scale (Figure 3A), we found highest measurement precision (i.e., lowest measurement error; see vertical axis) for low total scores but precision decreased as total score increased. For the depression scale (Figure 3B), highest precision was found for the lower total scores. For the restlessness scale, the total-score distribution was based on three items and heavily skewed, resulting in a valid estimate of the measurement precision for a restricted total-score range. Figure 3C shows that measurement precision was nearly constant at all scale levels.

Observed score distribution (connected open dots) and conditional error of measurement (solid line) for scales from Caci et al.’s (2003) three-factor model: (A) anxiety scale, (B) depression scale, (C) restlessness scale
Combined Results for Exploratory and Confirmatory MSA
Results from exploratory MSA were more similar to the Caci et al. (2003) three-factor model than results from confirmatory MSA. Exploratory MSA revealed two strong scales and one medium scale. One strong scale was identical to the anxiety scale from the Caci et al. three-factor model. Except for Item D5, the other strong scale contained the other five items from the Caci et al. depression scale. The medium scale contained two items from the Caci et al. restlessness scale. Confirmatory MSA of the Caci et al. scales showed good fit of the MHM.
Discussion
In the current study, we subjected the HADS, a frequently used measure of anxiety and depression in research, to an MSA as a method for analyzing questionnaire data collected in psychological research and clinical assessments, and contrasted the results of this analysis with that of factor analysis in a sample of cardiac patients treated with PCI. Combining the results of exploratory and confirmatory factor analyses, and exploratory and confirmatory MSA, we conclude that the HADS can best be viewed of as consisting of three scales, covering the attributes of depression, anxiety, and restlessness. This is consistent with the Caci et al. (2003) three-factor model, and also the results of other studies (Denollet et al., 2007; Dunbar et al., 2000; Friedman et al., 2001; Martin et al., 2003). More precisely, the two 5-item scales derived by exploratory MSA were reliable and covered the two key attributes of anxiety and (anhedonic) depression, suggesting using total scores on each set of five items to assess a patient’s anxiety and depression levels. Exploratory MSA also revealed a third scale containing the items A4 (“I can sit and feel relax”) and D7 (“I can enjoy a good book . . .”). These items relate to restlessness (Caci et al., 2003) or relaxed affect (Denollet et al., 2007). The reliability seems to be too low for individual decision making but high enough for research comparing groups from different clinical populations.
For the depression and restlessness scales, the EFA results differed somewhat from exploratory MSA results. Unlike EFA, exploratory MSA neither clustered Item D5 together with the depression items, nor Item A6 (“I feel restless as if I have to be on the move”) together with the restlessness items A4 and D7. These differences can be attributed to the restrictions with respect to lower bound c that MSA imposes on the item selection (here, c at least equal to .4). Both EFA and MSA found Items A6 and D5 to be weak indicators of the attribute, but only MSA imposed restrictions that rendered these items unscalable in the AISP. The AISP selects items one by one, and once an item is selected into a cluster it remains there. Such consecutive algorithms have the property that they sometimes lead to suboptimal scales. This explains why it was possible that confirmatory MSA found that the three items A4, A6, and D7 together constituted a medium restlessness scale, and amplifies the need for using several statistical methods rather than one in an effort to reduce method bias in the conclusions.
In their original article, Zigmond and Snaith (1983) presented the HADS as a collection of eight mandatory items (i.e., A1, A2, A3, and A5 for anxiety; D1, D2, D3, and D6 for depression), and six additional items. The mandatory items are included in the two 5-item scales found in the current study. Recently, Denollet et al. (2007) reduced the HADS to two 4-item scales: one scale including items D1, D2, D3, and D6, which according to Denollet et al.’s terminology measures reduced positive affect, and one scale consisting of items A1, A2, A3, and A7, which measures negative affect. Reduced positive affect and negative affect were independent predictors of major adverse clinical events in cardiac patients (Denollet et al., 2007).
We found that Item D5 (“lost appearance”) is a weak indicator of depression. Exploratory MSA did not cluster this item into the depression scale or any of the other scales. Confirmatory MSA showed that the item had low scalability with the other items, and CFA showed that the item had relatively low loadings on its corresponding factor. This favors the exclusion of Item D5 from the depression scale, corroborating the findings of Caci et al. (2003). Removal of items that have low scalability with other items is a simple way of obtaining conceptually clear scales but also means that less information is obtained for each individual and that reliability of individual decision making may be compromised (Emons, Sijtsma, & Meijer, 2007). Furthermore, removing items with low scalability may impair attribute coverage and hence threaten construct validity (e.g., Reise & Waller, 2009).
Alternatively, one could ask whether low scalability reflects that the item is a weak indicator of the attribute or that the item wording is poor. The latter problem may be solved using better wording, which may produce a strong indicator of the attribute. For example, the response options of Item D5 may show some ambiguity (“I have lost my interest in my appearance”): the first option is “definitely,” and refers to the extent to which one has lost interest in ones appearance, but another option (“I don’t take as much care as I should”) refers to the frequency with which attention is paid to one’s appearance. Thus, for one respondent the item refers to intensity and for the other it refers to frequency, and this ambiguity threatens the item’s validity.
Smith et al. (2006) used parametric IRT models to analyze the HADS in a large sample of cancer patients (see also Pallant & Tennant, 2007). Based on fit indices for the rating scale model, they concluded that the HADS has a higher order single factor structure with two unidimensional subscales. A subsequent Rasch model analysis of the depression and anxiety subscales showed misfit for Items A6, D5, and D7. Our confirmatory MSA showed that these items had much lower scalability than the other items (Table 5, columns 3 and 4), but we did not find significant violations of monotonicity. Hence, we conclude that the MHM fitted the items, implying ordinal measurement with relatively small contributions of Items A6, D5, and D7 to reliable person ordering. The Rasch model does not tolerate such aberrancies and requires items to be of the same quality, resulting in the removal of such items from the scale. The strictness of the Rasch model relative to the MHM explains the differences between Smith et al.’s (2006) results and ours.
With respect to the use of the HADS in research and clinical practice, our findings suggest that it may be reduced to a 10-item scale if one is only interested in measuring symptoms of anxiety and depression. The two 5-item scales tapping anxiety and depressive symptoms are structurally sound, and both subscale scores have an acceptable reliability. Combining the two scales results in a 10-item scale for measuring a higher order attribute for which H = .47 and alpha = .87. The total scores, however, may no longer have a clear interpretation. The separate anxiety and depression total scores correlated .57, indicating that patients with the same total score on the 10-item scale likely have different compositions of anxiety and depression symptoms. Thus, based on the composite total score alone one cannot tell whether the patient suffers from mostly depression, anxiety, or a mixture of the two. From a diagnostic point of view this is undesirable.
We limited this study to the analysis of the dimensionality and the measurement properties of the HADS, and offered some suggestions for improving the interpretation of HADS scores. Studies on the sensitivity and the specificity must be done to justify shortening the HADS to a 10-item version and to decide whether to use total scores or subscale scores for clinical diagnosis, so as to ensure that the screening properties of the scale for identifying patients with probable clinical anxiety and depression are preserved. This is particularly important, since an international expert committee has recommended the HADS as one of the screening tools for use in clinical cardiology practice to identify high-risk patients (Albus, Jordan, & Herrmann-Lingen, 2004). Smith et al. (2006) investigated the screening efficacy of the total scale and the subscale scores of the HADS and found that removal of items misfitting under the Rasch model had little impact on the screening efficacy of either scale. Our confirmatory MSA identified these items as weakly scalable. Unfortunately, we were not able to examine the screening properties of a 10-item version of the HADS, as we did not administer a clinical, diagnostic interview to assess anxiety and depressive disorder in the current study.
If future studies show that a 10-item version of the HADS provides an optimal balance between sensitivity and specificity, then discarding the restlessness subscale would result in a simple scoring algorithm for use of the HADS in clinical practice (Martin et al., 2008). Moreover, Martin et al. (2008) noted that the inclusion of the items loading highest on the third factor may contribute to false-positive and false-negative cases, if the full HADS is used as a screening instrument. They suggested that health care providers must be aware of this risk, but we suggest that a more pragmatic approach may be to exclude these items altogether if they do not contribute to the validity of the scale.
We conclude that our study supports the notion that the HADS has an underlying three-factor structure. However, since the reliability of the 3-item restlessness subscale is below the acceptable norm for individual diagnosis (Nunnally & Bernstein, 1994, p. 265), the HADS might best be used as a 10-item scale, containing two structurally valid scales for anxiety and depression. Given cumulative evidence that the HADS is composed of three factors rather than the two factors originally proposed by Zigmond and Snaith (1983), studies are now warranted that address the sensitivity and specificity of the 10-item measure proposed here, prior to considering adopting this version in research and clinical practice.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interests with respect to the authorship and/or publication of this article.
Funding
The author(s) received no financial support for the research and/or authorship of this article.
