Abstract
The Toronto Alexithymia Scale–20 is arguably the most utilized measure of alexithymia. Although a three-factor solution has been found by numerous studies, these findings are not universal. This article examined and compared 18 competing factor structures for the Toronto Alexithymia Scale–20, which included between one and four correlated latent factor structures, common methods models that accounts for negatively worded items, and bifactor models. Although the two-factor bifactor model with a common methods factor had the better model fit compared with the other 17 models examined, it still did not achieve the requisites of a good model fit across all model fit indices. Issues stemmed primarily from the externally oriented thinking factor and the negatively worded items. Post hoc analyses indicated that a two-factor bifactor model with the negatively worded items dropped achieved the requisites of a good model fit and can be treated as a unidimensional measure despite the presence of multidimensionality. Multiple-group analysis indicated that the factor loadings were invariant across U.S. and Philippines samples. After controlling for noninvariance at the item intercept level, the Philippines sample had a higher alexithymia general score compared with the U.S. sample.
Keywords
The construct alexithymia was initially coined in the 1970s to characterize individuals with difficulties identifying, processing, and describing emotions (Nemiah, Freyberger, & Sifneos, 1976). Arguably, the most utilized and most cited measure for alexithymia is the Toronto Alexithymia Scale–20 items (TAS-20; Bagby, Parker, & Taylor, 1994; Bagby, Taylor, & Parker, 1994), a measure that has been evaluated in numerous studies using student, community, and patient samples, and has been translated and evaluated in numerous cultures and countries (for reviews, see Kooiman, Spinhoven, & Trijsburg, 2002; Taylor, Bagby, & Parker, 2003). The TAS-20 has a three-factor structure: Difficulty Identifying Feelings (DIF; 7 items; e.g., “I am often confused about what emotion I am feeling”), Difficulty Describing Feelings (DDF; 5 items; e.g., “I find it hard to describe how I feel about people”), and Externally Oriented Thinking (EOT; 8 items; e.g., “I prefer to analyze problems rather than just describe them”). Although the DIF-DDF-EOT three-factor solution has been replicated numerous times (e.g., Loas et al., 2001; Meganck, Vanheule, & Desmet, 2008; Parker, Taylor, & Bagby, 2003; Preece, Becerra, Robinson, & Dandy, 2018; Taylor et al., 2003; Tsaousis et al., 2010), several studies have found alternative structures ranging from one- to four-factor solutions, with samples coming from both Western and non-Western countries (e.g., Cleland, Magura, Foote, Rosenblum, & Kosanke, 2005; Erni, Lötscher & Modestin, 1997; Haviland & Reise, 1996; Lambert et al., 1999; Zhu et al., 2007). Other psychometric issues have also been pointed out, including low model fit indices and poor reliability estimates for the EOT factor (Gignac, Palmer, & Stough, 2007).
The current study continues the long line of research on the psychometric properties of the TAS-20, particularly its factor structure and reliability. This study, however, expands the literature by examining 18 competing latent factor structures, including a correlated latent factor (CLF) model, common methods model, and bifactor model. Furthermore, the current study examined TAS-20 measurement and structural invariance between the United States and the Philippines, a country in which TAS-20 psychometric properties have yet to be studied.
Competing Factor Solutions of the TAS-20
Examination of a measurement’s factor structure is essential for theoretical reasons, particularly in ascertaining construct validity and accurate specifications of theory (Brown, 2014; Smith & McCarthy, 1995). Failure to ascertain the correct factor structure could lead to inaccurate correlational and experimental findings. Accuracy in the interpretation of scores is essential especially when an instrument is used for clinical purposes. Factor analytic studies can help determine how an instrument should be scored. That is, depending on the results, an instrument can be aggregated or summed if evidence of unidimensionality is present. Otherwise, subscale scores are used especially when multidimensionality is found. In this section, competing factor solutions for the TAS-20 are discussed.
Competing Correlated Latent Factor Models
The three-factor solution (DIF-DDF-EOT) was initially proposed by Bagby, Parker, et al. (1994) and was derived using factor analysis, specifically principal axis factoring. Some have used exploratory factor analysis or principal components analysis (e.g., Kojima, Frasure-Smith, & Lespérance, 2001), but more recent psychometric examinations of the TAS-20 have utilized confirmatory factor analysis (CFA) to ascertain its factor structure (e.g., Parker et al., 2003; Tsaousis et al., 2010). Studies using CFA typically utilize a CLF model which suggests that items uniquely load onto latent factors, which in turn are correlated with each other. Although a large number of studies from different countries and cultures utilizing student, adult, and inpatient populations have indicated the adequacy of the three-factor DIF-DDF-EOT solution (for reviews, see, e.g., Kooiman et al., 2002; Parker et al., 2003), this finding is not universal. Lambert et al. (1999) proposed a one-factor solution. Others suggested a two-factor solution wherein the DIF and the DDF were aggregated, leaving a DI/DDF and EOT factor structure (e.g., Cleland et al., 2005; Erni et al., 1997). An alternative three-factor solution was proposed by Haviland and Reise (1996) with a DI/DDF latent factor and splitting the EOT factor into pragmatic thinking (PT) and lack of importance of emotions (IM; hereinafter referred to as DI/DDF-PT-IM model). Finally, Müller, Bühner, and Ellgring (2003) proposed a four-factor DIF-DDF-PT-IM solution to the TAS-20. Gignac et al. (2007) have also questioned prior research that validates the TAS-20 DIF-DDF-EOT factor structure. Specifically, Gignac et al. (2007) pointed out that only 3 out of 27 studies mentioned in a review article (Taylor et al., 2003) achieved a goodness-of-fit index >.94 (for a rebuttal, see Bagby, Taylor, Quilty, & Parker, 2007).
Common Methods Model
Another issue with the TAS-20 DIF-DDF-EOT factor structure is the poor reliability estimates for EOT consistently found among multiple studies. For instance, in reviewing the TAS-20 psychometric properties across 22 countries (Taylor et al., 2003), the average Cronbach’s α of EOT is .57 (range: .27-.83). Some have hypothesized that the poor reliability could be because four out of five negatively worded items in the TAS-20 are in EOT (Gignac et al., 2007; Kojima et al., 2001; Preece et al., 2018). Within a CFA framework, this common methods bias issue (Podsakoff, MacKenzie, & Podsakoff, 2012) can be empirically examined and alleviated by adding an orthogonal latent factor that accounts for the negatively worded items. The orthogonal latent methods factor accounts for the variance shared by the items sharing a common method (in this case, negatively worded items) but are presumed not to be associated with the latent factors accounting for the traits. Studies that examined methods bias of negatively keyed items of TAS-20 suggest model improvement over a model without a common methods factor (Gignac et al., 2007; Mattila et al., 2010; Moriguchi et al., 2007; Preece et al., 2018; Watters, Taylor, Ayearst, & Bagby, 2016).
Bifactor Model Solution to the TAS-20
A common practice with TAS-20 is to use the summed score, either treating the variable as a continuum or utilizing a categorical approach where scores ≥61 indicate alexithymia (e.g., Dehgani, Dehgani, Kafaie, & Taghizadeh, 2017; Parker, Taylor, & Bagby, 1998). However, summing scores and treating the TAS-20 as a unidimensional measure disregards the multidimensionality consistently found in psychometric studies. There is, therefore, a tension regarding whether to conceptualize and apply TAS-20 as a unidimensional or a multidimensional measure. One source for this tension stems from psychometric studies that utilize a CLF perspective in which a one-factor solution is pitted against a multifactor solution, and studies have consistently shown the superiority of a multidimensional solution (e.g., Müller et al., 2003; Zhu et al., 2007).
One way to reconcile the unidimensionality versus multidimensionality argument is through a bifactor model rather than using a CLF. A bifactor structural model presumes that relationships among the items can be accounted for by a single general factor (in this case, alexithymia) and group factors account for additional variance common among the items, typically due to similarity in content (Reise, 2012). Using a bifactor model also has the advantage of evaluating whether the TAS-20 can be treated as a unidimensional construct, or whether the multidimensionality is so severe as to preclude the use of summed scores. To our knowledge, only two studies have examined a bifactor solution to TAS-20 (Gignac et al., 2007; Reise, Bonifay, & Haviland, 2013), and both have shown improvements in model fit in the bifactor model compared with the CLF model. Furthermore, support for a bifactor solution was found for a similar alexithymia instrument, the Toronto Structured Interview for Alexithymia (Watters, Taylor, & Bagby, 2016), which lends credence to a unidimensional conceptualization of alexithymia.
Cross-Country Differences in TAS-20 Factor Structures
Although the DIF-DDF-EOT factor structure has been replicated in various countries (Taylor et al., 2003), there are exemptions to this finding. For instance, in the Chinese translation of the TAS-20, a four-factor structure showed better fit compared with a three-factor model in an undergraduate sample, and Chinese student samples had significantly higher TAS-20 scores compared with a Canadian student sample (Zhu et al., 2007). In a Dutch student and patient sample, a two-factor solution was a better solution compared with the DIF-DDF-EOT (Kooiman et al., 2002). Given these discrepancies in the literature, there is a need to further examine the TAS-20 factor structure in various countries and cultures, particularly in non-Western countries, to avoid construct bias (Van de Vijver & Leung, 1997). Construct bias occurs when the construct purportedly being measured is not identical, is not present, or is not defined similarly in another culture (Van de Vijver & Leung, 1997). At a theoretical level, it is essential to assure that the alexithymia construct and its subfactors do exist and are defined and structured similarly in other cultures to establish cross-cultural generalizability and comparability.
Related to the construct bias issue, there is an extensive literature suggesting that culture has an impact on how emotions are appraised, how intense they are experienced, and how people react to and regulate emotions (e.g., Kitayama, Markus, & Matsumoto, 1995, Matsumoto et al., 2008; Schimmack, Oishi, & Diener, 2002). For instance, Western or individualistic cultures find it more normative to express emotions compared to Eastern or collectivist cultures (Matsumoto et al., 2008). Furthermore, Asian cultures, compared with non-Asian cultures, do not perceive emotions of opposite valence (e.g., happy and sad) as necessarily opposite, but rather compatible with each other (Schimmack et al., 2002). The West/individualist versus East/collectivist difference in emotional expression and regulation further highlights the need to examine alexithymia’s factor structure and scores cross-culturally. In this article, we examined two cultures that embody the West/individualist and East/collectivist difference: the United Sates and the Philippines.
An East–West country difference in TAS-20 scores has been documented. For instance, using the statistics reported by Taylor et al. (2003), Japanese students (M = 53.20, SD = 12.10, n = 473) had significantly higher alexithymia scores compared to a Dutch student sample (M = 43.93, SD = 9.12, n = 414; t = 12.74, degrees of freedom (df) = 885, p < .01). In an aggregated Arab-speaking sample from Algeria, Gaza, and Oman, alexithymia total scores and DIF, DDF, and EOT scores were significantly higher compared with a Canadian sample (El Abiddine et al., 2017). One limitation of simple sample comparisons using analysis of variance (ANOVA)–based procedures is that it is unclear whether the sample differences are due to real differences at the latent factor level (or “true score”) or due to bias or systemic variability at the item level (i.e., item differences at equivalent levels of the latent factor score or “true score”; Brown, 2014). One way to alleviate this concern is through the use of multiple groups analysis (MG-CFA; Chen, 2008; Meredith, 1993; Vandenberg & Lance, 2000), a procedure that can parse out whether group differences lie in the latent factor (structural model) or at the item factor loading or intercept level (measurement model; Brown, 2014).
Study Overview and Proposed Model Comparisons
Although the three-factor DIF-DDF-EOT (Bagby, Parker, et al., 1994) has received the most support, other studies have presented alternative factor structures. This article aims to help clarify the inconsistent findings regarding the factor structure of the TAS-20. Previous research has suggested five competing models: (a) a one-factor model (ALEX model, see Figure 1a), (b) a two-factor model (DI/DDF-EOT model, see Figure 1b), (c) a three-factor model proposed by Bagby, Parker, et al. (1994; DIF-DDF-EOT model, see Figure 1c), (d) a three-factor model proposed by Haviland and Reise (1996; DI/DDF-PT-IM model, see Figure 1d), and (e) a four-factor model (Müller et al., 2003; DIF-DDF-PT-IM model, see Figure 1e). These models assume that each item uniquely loads to a specific latent factor, and all latent factors are correlated with each other. These models are referred to as CLF models.

Measurement model of the one- to four-factor correlated latent factor model.
Some studies have also raised the issue posed by negatively worded items, which, as suggested, could affect the model fit indices. One way to address this issue from a CFA framework is to add an orthogonal latent factor that specifically accounts for negatively worded items. Figure 2 presents a common method factor added to the DIF-DDF-EOT model. The current study examined the viability of adding a common method factor to all five models (i.e., ALEX, DI/DDF-EOT, DIF-DDF-EOT, DI/DDF-PT-IM, DIF-DDF-PT-IM), referred to as common methods (CM) models. For example, the three-factor Bagby, Parker, et al. (1994) model with an orthogonal common method latent factor is referred to as DIF-DDF-EOT+CM model (see Figure 2).

Measurement model of the DIF-DDF-EOT common method factor model.
A bifactor model (herein referred to as BF models) presumes that a general factor (e.g., alexithymia) accounts for the relationship among the items, with group or specific factors accounting for the additional variance among the items, typically due to similarities in content. In other words, a BF model assumes that the TAS-20 measures a general alexithymia construct, and additional variance is accounted for by similarities in content (e.g., for DIF, all items pertain to identification of emotion). Procedurally, a BF CFA would have (a) one latent factor that accounts for the general factor (alexithymia) and all 20 items loading onto the general factor; (b) between two to four latent factors that account for similarities in content (in this case, we used the DI/DDF-EOT, DIF-DDF-EOT, DI/DDF-PT-IM, and DIF-DDF-PT-IM as group factors); and (c) all latent factors are orthogonal. Figure 3 presents the structural model of the DIF-DDF-EOT model with an added BF latent factor and is referred to as the DIF-DDF-EOT+BF model.

Measurement model of the DIF-DDF-EOT bifactor model.
It is also possible to combine a BF and a CM into one model, with a general alexithymia latent factor, a common methods latent factor to account for negatively worded items, and between two to four latent group factors to account for similarities in item content. Figure 4 presents the structural model of the DIF-DDF-EOT model with an added BF and CM latent factors, and is referred to as the DIF-DDF-EOT+CM+BF model.

Measurement model of the DIF-DDF-EOT bifactor and common methods factor model.
In summary, this study first aims to examine and compare 18 competing TAS-20 factor structure models using both nested and nonnested model comparisons. Specifically, the 18 models include five models that utilize CLF (i.e., ALEX, DI/DDF-EOT, DIF-DDF-EOT, DI/DDF-PT-IM, DIF-DDF-PT-IM), five models that expand the CLF models by adding a common methods factor (CM models), four BF models that will use the DI/DDF-EOT, DIF-DDF-EOT, DI/DDF-PT-IM, DIF-DDF-PT-IM as group or specific factors and four models that combine BF and CM models. Given the cross-country differences found in the literature, a secondary aim of the current study was to examine the TAS-20 factor structures in the United States and the Philippines side by side and compare these using MG-CFA.
Method
Participants and Procedures
Data for this study were from a larger study on sexual aggression perpetration and victimization between U.S. and Philippines samples. For the U.S. sample, 1,621 undergraduate students (74% female; age M = 19.83, SD = 2.62, range 17-57 years) were recruited from a large public plains state university and a private plains state university. The ethnic composition of U.S. participants was as follows: European American (n = 1,286, 79%), Asian Americans/Pacific Islander (n = 128, 8%), Hispanic (n = 78, 5%), African American (n = 49, 3%), Native American (n = 12, 1%), and other/rather not report (n = 68, 4%). For the Philippines sample, 482 undergraduate students (76% female; age M = 17.77, SD = 1.39, range 16-28 years) were recruited from a large public university from the Visayas region. Participants answered all the measures online and received course credits. All measures were in English. Prior to participant recruitment, the University of Nebraska–Lincoln and Creighton University Institutional Review Board reviewed and approved the research protocol. Institutional review boards are not a common practice in the Philippines; however, the University of the Philippines–Visayas Office of the Vice Chancellor for Research and Extension approved the protocol for this study. All participants were presented with and signed informed consents prior to participating in the study.
Measurements
Toronto Alexithymia Scale–20 (TAS-20)
The TAS-20 (Bagby, Parker, et al., 1994) is a self-report instrument that measures alexithymia. Participants indicated their agreement of each item using a Likert-type scale ranging from 1 (strongly disagree) to 5 (strongly agree). Total scores can range from 20 to 100, with scores ≥61 indicative of alexithymia, 51 to 60 as borderline, and ≤50 as no alexithymia. As previously noted, the TAS-20 is commonly reported to have a three-factor DIF-DDF-EOT solution (Taylor et al., 2003); however, this finding is not universal. Previous research reported α reliability coefficient ranges from .68 to .84 for the total score, from .67 to .85 for DIF, from .48 to .82 for the DDF, and from .27 to .83 for the EOT factor (Taylor et al., 2003). Reliability estimates for the current study are reported in the Results section.
Data Analysis
CFA, Nested and Nonnested Model Comparisons
CFA on the 18 alternative models to the TAS-20 was performed using Mplus version 6 (Muthén & Muthén, 2010). For estimation procedure, maximum likelihood with robust standard errors was utilized to account for multivariate nonnormality and missing data (full-information maximum likelihood). Adequacy of model fit for each model was evaluated using the comparative fit index (CFI), Tucker–Lewis index (TLI), root mean square error of approximation (RMSEA), and the standardized root mean square residual (SRMR). Hu and Bentler (1999) suggested the following criteria for a good model fit: CFI ≥ .95, TLI ≥ .95, RMSEA ≤ .06, and SRMR ≤. 08. Brown (2014) suggested a lower criterion of CFI ≥ .90 and TLI ≥ .90. For this study, the following criteria was used to assess a good model fit: CFI ≥ .90, TLI ≥ .90, RMSEA ≤ .06, and SRMR ≤ .08.
Nested and nonnested model comparisons were performed to compare various TAS-20 factor models. Nested model comparisons are conventionally evaluated using a Δχ2 test; however, because of the use of MLR, a likelihood ratio test accounting for the scaling correction factors was used to make nested model comparisons (−2ΔLLcorrected; Satorra, 2000). Similar to the Δχ2 test, the −2ΔLLcorrected is evaluated using a χ2 distribution, and a p < .05 suggests that the model with an additional parameter (e.g., adding a latent factor) provided a better fit to the data compared with a more parsimonious one. For nonnested model comparisons, the Akaike information criterion and the Bayesian information criterion (BIC) is generally used, with lower values indicating better model fit (Brown, 2014; Burnham & Anderson, 2004). Merkle, You, and Preacher (2016) though have raised concerns regarding the use of “lower is better” criterion and advocated for more stringent procedures such as the Vuong (1989) test for nonnested model comparisons. Vuong (1989) test utilizes a z distribution, with p < .05 indicating a significant difference in BICs between two models, and a smaller BIC suggesting a better model fit.
Bifactor Model Statistics
Statistics are available to evaluate the uni- or multidimensionality of a bifactor model and to assess how well the general factor accounts for the variance among the items. The alpha (α) and omega (ω) reliability estimates are reported in this study. For bifactor models, the omega reliability (ω) measures both the variance accounted for by the general alexithymia factor and the other specific factors. On the other hand, the omega hierarchical (ω H ) reflects the total score variance attributable to the general alexithymia factor only after accounting for all specific factors, with values greater than .75 preferred (Reise, Scheines, Widaman, & Haviland, 2013). In other words, a high ω and a high ω H indicate that the general factor accounts for most of the variance in the model. The omega hierarchical subscale (ω HS ) reflects the subscale score variance after accounting for the general factor. Low ω HS (i.e., less than .50; Reise, Bonifay, et al., 2013) suggests that majority of the subscale score variance is due to the general factor, with the leftover variances accounted for by the specific factors (e.g., similarities in the items not otherwise accounted for by the general factor).
The explained common variance of the general factor (ECVGen) is an index of unidimensionality, which could be interpreted as the relative strength of the general factor versus specific factors (Reise, Moore, & Haviland, 2010). An ECVGen ≥.85 indicates little common variance after accounting for the general factor, enough to consider the measure as unidimensional (Reise et al., 2010). To further assess for unidimensionality, the ECVGen and the ω H are interpreted with the percent of uncontaminated correlations (PUC), which is the number of uncontaminated correlations divided by the number of unique correlations (see Rodriguez, Reise, & Haviland, 2016, for PUC equation). According to Reise, Scheines, et al. (2013), when PUC < .80, an instrument with ECVGen > .60 and ω H >.70 could be treated as unidimensional despite the presence of common or specific factors, or, for purposes of this article, “unidimensional enough” (Reise et al., 2013).
Multiple Groups CFA (MG-CFA)
Another aim for this article was to compare TAS-20 factor structures between U.S. and Philippines samples. To achieve this goal, CFA was performed for both samples independently. Once a factor structure solution with the better model fit was decided on, a MG-CFA (Chen, 2008; Meredith, 1993; Vandenberg & Lance, 2000) was subsequently performed to evaluate the measurement and structural invariance of the factor structure between samples.
MG-CFA measurement invariance starts with estimating the configural model, that is, a model wherein all parameters are allowed to vary across the U.S. and Philippines samples. A configural model is needed to assure that similar factors are measured in each group, and to establish a baseline for which more restrictive models can be compared. All subsequent models are examined with more restrictions. A metric invariance model is subsequently calculated by constraining unstandardized factor loadings to be equal between samples. A significant −2ΔLLcorrected test performed between the configural model and metric invariance model indicates possible nonequivalence at the factor loading level, and sources of model misfit are examined using the MODINDICES. Cheung and Rensvold (2002) also suggests that a ΔCFI > .01 suggests that the assumption of noninvariance should be rejected. Metric invariance assures that the strength of the items-factor relationship is similar across samples.
Item intercepts are subsequently constrained to be equal across samples (scalar invariance model), and model fit is compared against the metric invariance model using procedures previously discussed. The scalar invariance test examines whether item means are proportionally equal across groups. The residual variance invariance model (for residual variance) and residual covariance invariance model (for residual covariance) are then evaluated following the procedures previously outlined.
After evaluating the measurement invariance model, the structural model invariance models are subsequently examined. Structural invariance tests whether the factor means and the relationship between latent factors are equal across samples. The latent factor variance invariance was first examined, followed by latent factor covariance, and finally the latent factor means. Structural invariance tests were evaluated using the −2ΔLLcorrected and ΔCFI test.
Results
Examination of the Model Fit Indices
In efforts to evaluate and compare various TAS-20 factor structures, none of the models examined achieved the minimum criteria for the TLI (i.e., >.90; see Table 1 for model fit indices). Excluding TLI, only seven models achieved the minimum requisites for the CFI, RMSEA, and SRMR among all the models examined. For the U.S. sample, these include the DI/DDF-EOT+CM+BF, DIF-DDF-EOT+CM+BF, DI/DDF-PT-IM+CM+BF, and the DIF-DDF-PT-IM+CM+BF models. However, the DI/DDF-PT-IM+CM+BF and the DIF-DDF-PT-IM+CM+BF models resulted in a negative residual variance at Item 15; hence, these models were not considered viable. For the Philippines sample, the DIF-DDF-PT-IM+CM, DI/DDF-EOT+CM+BF, and the DIF-DDF-EOT+CM+BF models achieved minimum criteria of good model fit for the CFI, RMSEA, and the SRMR. However, the DIF-DDF-PT-IM+CM model resulted in latent factor correlations greater than 1 in IM and PT; hence, these were not considered viable models.
Model Fit Indices for the Correlated Latent Factor Model, Common Methods Factors Model, and Bifactor Model.
Note. CFI = comparative fit index; TLI = Tucker–Lewis index; df = degrees of freedom; RMSEA = root mean square error of approximation; pclosefit = RMSEA p of close fit; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion; ALEX = alexithymia; DI/DDF = Difficulty Identifying/Difficulty Describing Feelings; DDF = Difficulty Describing Feelings; DIF = Difficulty Identifying Feelings; EOT = Externally Oriented Thinking; IM = Importance of Emotions; PT = Pragmatic Thinking. Figures in bold indicate exceeding minimum requisite of a good model fit for the specific model fit index.
Resulted in a linear dependency between IM and PT latent factors, with latent factor correlations >1. bNo Convergence, number of iterations exceeded. Issue stems from a negative residual estimate in Item 5. cNo convergence, number of iterations exceeded. Issue stems from a negative residual estimate in Item 8. dModel resulted in a negative residual variance in Item 15 for the three-factor DI/DDF-PT-IM+CM+BF (−2.184) and the four-factor DIF-DDF-PT-IM+CM+BF model (−1.211).
p < .01.
We subjected the viable models (i.e., DI/DDF-EOT+CM+BF, DIF-DDF-EOT+CM+BF) to model comparisons to ascertain which is a better model. Because all viable factor structure models are nonnested, a Vuong z test was utilized to make model comparisons. For the U.S. sample, although the BIC for the DIF-DDF-EOT+CM+BF model was smaller than DI/DDF-EOT+CM+BF, the difference was not significant (Vuong z = 0.379, p = .65). For the Philippines sample, the DI/DDF-EOT+CM+BF model BIC was smaller than the DIF-DDF-EOT+CM+BF; however, the difference was not significant (Vuong z = −0.578, p = .28). Given that the DI/DDF-EOT+CM+BF and the DIF-DDF-EOT+CM+BF for both samples were equivalent, we opted for a more parsimonious DI/DDF-EOT+CM+BF model. In the other CLF and CM analyses conducted in this study, latent factor correlations between the DDF and DIF were high for the correlated factor model (U.S. r = .835, p < .01; Philippines r = .935, p < .01) and the common factor model (U.S. r = .827, p < .01; Philippines r = .936, p < .01), further supporting the idea that the DDF and DIF might be measuring similar constructs. A more comprehensive nested and nonnested pairwise model fit comparison is available in the online Supplementary Material (Table S1 and S2).
Examining Factor Loadings, Reliability, and Unidimensionality
Standardized factor loadings for other models examined are available in the Supplementary Material (Table S3 to Table S10) but, due to page limitations, are not reported here. Table 2 presents the standardized factor loadings for the DI/DDF-EOT+CM+BF model, the model that best fit among the 18 models, for the U.S. and Philippines samples. Several observations need to be highlighted. For the U.S. sample, negatively worded items (Items 5, 10, 18, and 19) in the EOT factor either did not significantly load or had poor loadings (i.e., <.30) in the EOT latent factor and the general alexithymia factor. For the Philippines sample, negatively worded items significantly loaded on the EOT latent factor, although did not load on the general alexithymia general factor or the common methods latent factor. This could indicate that the EOT latent factor is accounting for the negatively worded nature of the items rather than tapping externally oriented thinking. An examination of reliability estimates α and ω indicate good reliability for the total alexithymia scale and the DI/DDF subscale, but not for the EOT subscale. Poor reliability estimates for EOT were also observed for other latent factor solutions (see also Supplementary Material Table S2 and Table S3 for reliability estimates for other factor solutions).
Standardized Parameter Estimates for the DI/DDF+CM+BF Models: U.S. and Philippines Samples.
Note. DI/DDF = Difficulty Identifying/Difficulty Describing Feelings; EOT = Externally Oriented Thinking; BF = Alexithymia general factor; CM = common methods factor—latent factor for negatively worded items. ECV = explained common variance. ω H = omega hierarchical; ω HS = omega hierarchical subscale. Percent of uncontaminated correlations for this analysis is .452.
p < .05.
Reliability and bifactor indices such as α, ω, ω H , ω HS , and ECV are available in Table 2. For the DI/DDF-EOT+CM+BF model, the general alexithymia factor accounts for 35% (Philippines sample) to 57% (U.S. sample) of the common variance, whereas 65% to 43% of the common variance is accounted for by the other group factors (see Table 2, ECV estimates under the BF column). To further assess for unidimensionality, the ECVGen and the ω H is interpreted with the PUC. As stated above, an instrument with PUC < .80, ECVGen >.60 and ω H >.70 could be treated as unidimensional despite the presence of common or specific factors. As Table 2 indicates, the assumption of unidimensionality has not been attained for both the U.S. and Philippines samples.
Overall, the results here suggest that, although a bifactor with common methods factor has the better model fit compared with the other 17 models examined, unidimensionality cannot be assumed. Results question the practice of using the sum of scores for TAS-20 and treating it as a unidimensional measure. Results also suggest that negatively worded items loaded poorly, and when comparing the U.S. and Philippines samples, inconsistently. The reliability of EOT, the latent factor where the majority of the negatively worded items loaded, was consistently shown to be poor. These observations are consistent with other studies that examined the psychometric properties of the TAS-20 (Gignac et al., 2007). Given the issues presented by negatively worded items, we subsequently examined a bifactor solution with these negatively worded items dropped from the model. A bifactor solution was examined given our results indicating the superiority of a bifactor solution compared with other factor solutions, as well as those suggested by Gignac et al. (2007) and Reise et al. (2013).
Post Hoc Analyses: Dropping Negatively Worded Items
For the U.S. sample, results indicated that the model fit indices of the bifactor model with negatively worded items dropped achieved the requisites of a good model fit except for RMSEA (CFI = .932, TLI = .905, RMSEA = .063, SRMR = .038). Modification index indicated that adding residual covariance between Items 3 and 7 would significantly improve model fit. As both items pertain to physical or bodily sensations, there is a conceptual rationale to add this parameter. A TAS-20 bifactor model with residual covariance between Items 3 and 7 achieved the requisites of a good model fit for the U.S. and Philippines sample (See Table 1). Table 3 presents the standardized factor loadings for this model.
Standardized Parameter Estimates for the Bifactor Models With Negatively Worded Items Dropped.
Note. BF = Alexithymia general factor; DI/DD = Difficulty Identifying/Difficulty Describing Feelings; EOT = Externally Oriented Thinking. ECV = explained common variance. ω H = omega hierarchical; ω HS = omega hierarchical subscale. Percent of uncontaminated correlation for this analysis for this analysis is .419.
p < .05.
To review, a bifactor model with PUC <.80, ECVGen >.60, and ω H >.70 is enough to warrant a unidimensional interpretation despite the presence of multidimensionality stemming from specific factors. As presented in Table 3, these prerequisites have been achieved for the U.S. and Philippines samples. For the U.S. sample, with ω = .888, and ω H = .841, the majority of the reliable variance (percentage of reliable variance = .841/.888 = 95%; Rodriguez et al., 2015) is attributable to the general factor. Percentage of reliable variance for the Philippines sample is 95%. The average relative parameter bias (i.e., the difference between the item’s factor loading in a unidimensional solution and the factor loading of the general factor in the bifactor model, divided by the general factor loading) is 6% for both the U.S. and Philippines samples, well below the acceptable average parameter bias of 10% to 15% (Muthén, Kaplan, & Hollis, 1987).
Overall, results of this analysis indicate that dropping the negatively worded items and using a bifactor model on the TAS resulted in a solution with good model fit indices. In addition, this factor structure solution can be considered “unidimensional enough” for the summed scores to be used and the measure to be treated as a unidimensional measure of alexithymia. Despite these results, the EOT factor still suffers from poor reliability, and Items 15 and 16 loaded poorly on the general alexithymia factor and load heavily on the specific factor.
Multiple Groups Analysis
The secondary aim for this article was to evaluate the measurement and structural invariance of the best-fitting factor structure model for the TAS-20 between the U.S. and Philippines samples. Post hoc analysis indicated that the bifactor model with the negative items dropped achieved the requisites of a good model fit for both samples, and results of the MG-CFA for this model are presented here. Results of the MG-CFA for the DI/DDF-EOT+CM+BF model (best fitting model across the 18 competing factor structure models) are available in the Supplementary Material.
A configural invariance model was initially estimated, with the U.S. sample as the reference group. Metric invariance model was subsequently estimated by constraining unstandardized item factor loadings as equal across samples. The metric invariance model (weak invariance) did not result in a significant decrease in fit relative to the configural model (−2ΔLLcorrected = 34.850, Δdf = 27, p = .14, ΔCFI = −.004, ΔTLI = –.017). These results suggest that the same latent factor was being measured for the U.S. and the Philippines samples, and that the strength of the association between items and latent factors is relatively equal across groups.
The scalar model (strong invariance) was then estimated by constraining item intercepts equal across samples. Model comparison indicated a significant decrease in model fit compared with the metric invariance model (−2ΔLLcorrected = 87.744, Δdf = 12, p < .01, ΔCFI = .008, ΔTLI = .005). Examination of the modification indices suggested that allowing Items 1, 6, 7, and 13 to vary across samples would improve model fit. Intercepts for these items were allowed to freely vary between samples, and model comparisons suggested that the partial scalar model was not significantly worse compared with the metric model (−2ΔLLcorrected = 11.428, Δdf = 8, p = .179, ΔCFI = .001, ΔTLI = −.002). This indicates that the observed differences in item means between groups is due to factor mean differences, except for Items 1, 6, 7, and 13. Examination of item intercepts suggests that the U.S. sample, compared with the Philippines sample, had a lower item response for Item 1 (U.S. = 2.347, Philippines = 2.505), Item 6 (U.S. = 2.355, Philippines = 2.626), and Item 7 (U.S. = 1.933, Philippines = 2.167) but higher item response for Item 13 (U.S. = 2.053, Philippines = 1.918) at the same absolute trait level of the latent factors/constructs.
Due to the noninvariance in the scalar model, residual variances for Items 1, 6, 7, and 13 were allowed to vary across samples. Equality of the unstandardized residual variances across groups were subsequently examined (residual variance invariance model or strict invariance) by constraining residual variances to be equal across groups, except for Items 1, 6, 7, and 13. The residual variance model had a significantly worse model fit compared with the partial scalar model (−2ΔLLcorrected = 37.279, Δdf = 11, p < .01, ΔCFI = .004, ΔTLI = .001), and modification indices indicated that the model could be improved by allowing Items 3 and 14 residual variances to vary across groups (−2ΔLLcorrected = 12.848, Δdf = 9, p = .170, ΔCFI = .000, ΔTLI = −.002). Equality of residual covariance between Items 3 and 7 across groups was examined, and results indicated no significant decrease in model fit compared with the partial residual variance invariance model (−2ΔLLcorrected = 3.716, Δdf = 1, p = .054, ΔCFI = .001, ΔTLI = .000).
Structural invariance was tested between the U.S. and Philippines samples. Constraining latent factor variances to be equal across groups resulted in a worse model fit compared with the partial measurement model (−2ΔLLcorrected = 24.546, Δdf = 3, p < .01, ΔCFI = .002, ΔTLI = .002). Modification indices indicated that allowing latent factor variance for DI/DDF to vary across samples would not make the model significantly worse than the partial measurement model (−2ΔLLcorrected = 5.842, Δdf = 2, p = .059, ΔCFI = .000, ΔTLI = .000). With the U.S. sample as the reference group (DI/DDF latent factor variance = 1), the Philippines sample had less variability in the DI/DDF (latent factor variance = 0.516). Constraining latent factor means to be equal across samples resulted in a worse model fit compared to the partial factor variance invariance model (−2ΔLLcorrected = 256.523, Δdf = 3, p < .01, ΔCFI = .026, ΔTLI = .027), and modification indices suggest allowing the alexithymia general model to vary across samples. Following this suggestion resulted in a nonsignificant difference between this model and the partial factor variance invariance model (−2ΔLLcorrected = 1.770, Δdf = 2, p = .413, ΔCFI = .000, ΔTLI = .000). Compared with the U.S. as the reference group (BF alexithymia general latent factor mean = 0), the Philippines sample had a significantly higher level of general alexithymia (latent factor M = 0.996).
In summary, MG-CFA results suggest similar factor loadings between the U.S. and Philippines sample. There were group differences at the item intercept level, but even after controlling these item intercept differences, the Philippines sample had a higher alexithymia “true score” trait compared with the U.S. sample.
Discussion
TAS-20 Factor Structure
The majority of the psychometric studies on the TAS-20 indicate a three factor DIF-DDF-EOT structure (Taylor et al., 2003); however, this finding is not universal with several studies suggesting between one- to four-factor solutions, including CLF, CM, and BF formulations (e.g., Haviland & Reise, 1996; Müller et al., 2003). This study expanded the current literature by simultaneously comparing 18 competing TAS-20 latent factor solutions using nested and nonnested model comparisons in samples from the U.S. and the Philippines. Results of this study indicated that the DI/DDF-EOT+CM+BF model fit the data better than the other models examined. Several findings need to be emphasized. First, there was a high correlation between the DIF and DDF latent factors, whether it was a CLF or a CLF+CM model. This suggests that the DIF and the DDF factors might be measuring similar constructs or traits. This result is not surprising as both factors tend to tap deficiencies in emotional awareness, which necessarily involves difficulties in identifying and describing emotions. Aggregating the DIF and DDF into one factor leads to a more parsimonious model, and consequently, a better fitting model.
Second, poor α and ω reliability estimates were observed for the EOT, IM, and PT in all models examined. These results replicated previous studies (e.g., Kojima et al., 2001), and some researchers have attributed the poor reliability to negatively worded items (Gignac et al., 2007; Kojima et al., 2001). This study attempted to mitigate the impact of negatively worded items when the common methods models were examined. Third, when these common method models were examined, the negatively worded items had high factor loadings in the common method factor, and loaded very low in the EOT, IM, and PT factors. In other words, after the negatively worded nature of the items was accounted for, these items no longer loaded in the EOT, PT, or IM factors. The rationale for the practice of using positively and negatively worded items was initially to prevent response bias; however, this results in psychometric issues and poor reliability and model fit (Sonderen, Sanderman, & Coyne, 2013; see Weijters, Baumgartner, & Schillewaert, 2013, for issues presented by negatively worded items), which seems apparent in this study and in others (e.g., Gignac et al., 2007).
Fourth, across the 18 competing models examined, bifactor models (specifically, DI/DDF-EOT+CM+BF as the best fitting model) outperformed the CLF and CM models. This suggests that the TAS-20 is better accounted for by a general alexithymia construct, with subfactors accounting for methodological and content similarities among the items. A bifactor conceptualization of TAS-20 is consistent with other studies (Gignac et al., 2007; Reise et al. 2013), as well as the interview-based assessment of alexithymia (Toronto Structured Interview for Alexithymia; Watters, Taylor, & Bagby, 2016). Watters, Taylor, and Bagby (2016) even opined that a bifactor solution is more consistent with how alexithymia was originally conceptualized by Nemiah et al. (1976). The bifactor model, however, is not without criticisms. A central issue pertinent to this article is the tendency for “overfitting” when comparing bifactor models with other models, with bifactor models capturing unwanted noise (Bonifay, Lane, & Reise, 2016). Bonifay et al.’s (2016) solution to this problem is to use alternative statistics (i.e., PUC, ECV, ω, and ω H ) in conjunction with model fit statistics. As shown in Table 2, the DI/DDF-EOT+CM+BF’s bifactor alternative statistics are poor at best. It is only when the negatively worded items are dropped that bifactor model alternative statistics improved.
Because of the issues surrounding negatively worded items, we performed a post hoc bifactor CFA where said items were dropped. This model achieved the requisites of a good model fit, and other bifactor model statistics (PUC, ECV, ω H ) suggested a “unidimensional enough” model (Reise, Scheines, et al., 2013). This strategy, however, is not without its own set of issues. Despite dropping negatively worded items, the EOT factor still had poor reliability estimates. A closer inspection of the EOT items indicated low factor loadings on the EOT factor and the general alexithymia factor, which subsequently resulted in the poor reliability. This result could suggest that the EOT items, although measuring the alexithymia construct, measures it weakly. On the other hand, the weak factor loading of EOT items on the general alexithymia factor does not absolutely conclude that the EOT factor and related items are not good indicators of alexithymia. Dropping items runs the risk of ending up with construct underrepresentation; that is, the latent factor no longer measures what it initially intended to measure because core items have been dropped. Furthermore, because of the comparatively large number of DI/DDF items retained and only a handful of EOT remained, it can be argued that the general alexithymia factor purportedly being measured really taps into the difficulty identifying and describing emotions, but not the full spectrum of alexithymia.
Initial theoretical conceptualizations of alexithymia have always included EOT as a core part of alexithymia (Bagby, Parker, et al., 1994; Bagby, Taylor, et al., 1994). However, factor analytic studies are essential to validate whether EOT is indeed a core component of alexithymia, at least for the TAS-20. This article is far from resolving the issue given the methodological problems presented by the negatively worded items. One way to attain resolution is for future studies to replicate this study with the TAS-20 negatively worded items framed positively. If said replication study would still yield similar problems with the EOT, then the general conceptualization of alexithymia might need rethinking.
TAS-20 Factor Structure Between the United Stated and the Philippines
A secondary aim for this article was to examine the measurement and structural invariance of the TAS-20 between U.S. and Philippines samples. In the bifactor model with negatively worded items dropped, results indicated that the items were measuring similar latent factor constructs, specifically a general alexithymia construct. We also performed a MG-CFA to the DI/DDF-EOT+CM+BF model, and results indicated relative equivalence of the model between samples, at least at the factor loading level (see Supplementary Material). This indicates that the items in the model have cross cultural generalizability and applicability, particularly to the Philippines. In other words, the U.S. and Philippines samples both conceptualize alexithymia factors similarly and items tap the purported latent factor similarly across samples.
However, partial invariance was found at the item intercept level, and at the factor variance and factor mean level. The MG-CFA indicated that there were systematic differences at the item level, particularly Items 1, 6, 7, and 13. In other words, even if latent factor scores or “true scores” are set to be identical between the two groups, the items mentioned showed systemic bias at the intercept level. However, even after accounting for the group differences at the item intercept level (by allowing these item intercepts to vary across groups), factor mean invariance model results revealed that the Philippines sample had a higher alexithymia general factor score compared with the U.S. sample.
Literature on emotions in a Filipino sample is scarce, and for alexithymia, nonexistent. However, there is an extensive literature suggesting that culture has an impact on how emotions are appraised, how intense they are experienced, and how people react to and self-regulate emotions (e.g., Kitayama et al., 1995; Matsumoto et al., 2008). Among Filipinos, it is viable to posit that culture accounts for the differences in alexithymia scores. For instance, the emphasis on social desirability and interpersonal harmony and the primacy of the group over the individual (Agbayani-Siewert, 1994) could motivate Filipinos to suppress one’s emotions if expressing these emotions would interfere with the collective harmony. The Filipino cultural values of pakikisama (loosely translated as social acceptance or conformity) and hiya (loosely translated as shame) highlights this point. Pakikisama encourages Filipinos to remain in harmony with their peers rather than vocalize disagreements or express emotions that would threaten group cohesion (Nadal, 2011). Hiya dissuades Filipinos from bringing shame, disappointment, and embarrassment to the family, sometimes at the cost of suppressing one’s emotions and aggravating mental health issues (Nadal, 2011). More research on emotions and alexithymia among Filipinos is still needed however.
Clinical and Research Implications
Results of this study indicated that the assumption of unidimensionality was supported only for the model where the negative items were dropped (DI/DDF-EOT+BF). On the other hand, for the original TAS-20 (i.e., measure with negatively worded items included), the assumption of unidimensionality was not supported. These results have research and clinical implications. Using the total scores for the original TAS-20, whether for research or clinical purposes, is not recommended as specific factors, specifically the EOT, present multidimensionality issues. In contrast, if the negatively worded items were dropped, using summed scores is defensible, and conceptually, alexithymia can be considered as a unitary construct. The bifactor solution with negatively worded items dropped is not without problems though. Specifically, items in the EOT specific factor had either low standardized factor loadings on the alexithymia general factor compared to items in the DI/DDF specific factor (Items 8 and 20), or had a higher factor loading in the EOT specific factor and low factor loading on to alexithymia general factor (Items 15 and 16). At this point, it is unclear whether the issue is methodological (i.e., items were pertaining to “preferences”) or conceptual (i.e., whether EOT items are strong measures of alexithymia, or whether alexithymia is better measured with items pertaining to emotion identification).
This study documented partial invariance at the item intercept level and has shown factor mean differences between the U.S. and Philippines samples. Cross-country comparison studies using ANOVA-based statistics have documented differences in alexithymia and subscale scores (e.g., El Abiddine et al., 2017; Zhu et al., 2007). One limitation of ANOVA-based comparisons is that it is not known whether the differences lie in the “true score” or due to systemic bias at the item level. Results of this study indicated it can be both. Hence, it is important to first evaluate invariance at the item level before conducting cross-cultural comparisons in alexithymia. Furthermore, the use of cutoff scores (e.g., >61) in identifying individuals with alexithymia in different cultures is not encouraged as cultural factors and systemic item differences can affect TAS scores.
Limitations and Future Directions
The results of this study should be tempered by its limitations. First, having a college student sample limits the generalizability of the results, particularly given studies suggesting differences in factor structure between student and patient samples (e.g., Zhu et al., 2007). Future studies should replicate this study using a nonstudent sample. Furthermore, the U.S. sample is significantly older (t = 16.47, p < .01) and had more nontraditionally aged students compared with the Philippines sample, which could have accounted for the results and sample differences presented here. Second, the analyses performed and the conclusions reached in this study are predominantly statistically and methods-driven. For instance, the negatively worded items were dropped due to methodological rather than theoretical considerations, and said items could be essential to the conceptual definitions of alexithymia. Future research could examine the conceptual importance of the dropped items, as well as examine whether phrasing the negatively worded items positively would impact the factor structure solution and model fit. In addition, future research should further examine the psychometric properties and the validity of the TAS measure with negative items dropped. Finally, this study did not examine convergent and divergent validity, clinical utility, or the feasibility of a cutoff score, topics that future research can address.
Supplemental Material
Supplemental_Material – Supplemental material for Toronto Alexithymia Scale–20: Examining 18 Competing Factor Structure Solutions in a U.S. Sample and a Philippines Sample
Supplemental material, Supplemental_Material for Toronto Alexithymia Scale–20: Examining 18 Competing Factor Structure Solutions in a U.S. Sample and a Philippines Sample by Antover P. Tuliao, Alicia K. Klanecky, Bernice Vania N. Landoy and Dennis E. McChargue in Assessment
Footnotes
Acknowledgements
The authors would like to thank Dr. Jospeh H. Hammer and David Dueber for their assistance in the analysis.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
