Abstract
Despite the widespread use of RIASEC interest inventories, little is known about whether these inventories actually measure the same core constructs and provide similar career recommendations to individuals. This study investigates the construct validity among four major interest inventories—the Self-Directed Search (SDS), O*NET Interest Profiler (IP), ACT Interest Inventory (UNIACT), and Strong Interest Inventory (SII). Results showed that RIASEC interest scores from the four inventories were highly correlated, but the measures often gave respondents different high-point codes. Item content analysis revealed that the basic interests reflected in each RIASEC scale both overlapped and diverged across inventories, providing an explanation for why RIASEC inventories are not interchangeable. We integrate findings across our analyses to offer cautionary notes for choosing among established RIASEC inventories and interpreting interest results. Furthermore, we also provide recommendations for constructing the next generation of basic interest inventories.
Keywords
Interest inventories are widely used to help people make decisions in educational and occupational settings (Hanna & Rounds, 2020). Since the 1970s, Holland’s (1973, 1985, 1997) RIASEC model has been used to assess and organize interest types in vocational interest inventories. The RIASEC model provides a theoretical framework that helps individuals understand their interests in six general domains: Realistic (R), Investigative (I), Artistic (A), Social (S), Enterprising (E), and Conventional (C). Despite the popularity of Holland’s RIASEC model, few studies have compared what is being assessed by distinct inventories adopting the model. As a result, little is known about the convergent validity of RIASEC scores from different interest inventories.
Understanding the convergence and content coverage among interest inventories is critical for both practice and research. In practice, individuals (and career counselors) often use RIASEC interest profiles to help focus career and educational exploration. Occupations are typically considered good fits when they have similar patterns of RIASEC scores to individuals. To work effectively on a large scale, this method of career matching requires that interest inventories bear similar reflections of the RIASEC constructs. If interest inventories do not converge on their representation of Holland’s model, an individual could receive different RIASEC profiles from different inventories, and would therefore be pointed towards divergent careers. Convergent validity is also critical for research purposes. The RIASEC types are complex, multidimensional constructs that each cover a range of basic interests (Su et al., 2019). For cumulative knowledge to progress about the structure, development, and correlates of vocational interests, it is essential that researchers are aware of how item content both overlaps and diverges across different measures.
In this study, we investigate evidence for convergence and content validity among four of the most widely used interest inventories. We use both statistical and content-based approaches to examine: (a) the Self-Directed Search (SDS; Holland et al., 1994), (b) the O*NET Interest Profiler Short Form (IP; Rounds et al., 2021), (c) the ACT Interest Inventory (UNIACT-R; ACT, 1995), and (d) the Strong Interest Inventory (SII; Harmon et al., 1994). Overall, our study has three major aims. First, we examine interest score convergence by fitting Confirmatory Factor Analysis (CFA) models to the multitrait-multimethod (MTMM) matrix to reveal the relative influence of trait and method factors in influencing interest scores. Second, we build on the CFA-MTMM approach to assess convergence in terms of RIASEC high-point code agreements, a central issue for career counseling. Third, we systematically examine item content coverage using basic interests, which are narrower interest categories that underly the RIASEC themes (Su et al., 2019). Each approach provides distinct contributions to better understanding interest measurement—the first two quantitative approaches update current understanding of the convergent validity among RIASEC measures, whereas the content analysis provides critical insight into what is actually measured by different inventories. Based on our results, we offer guidance to both practitioners and researchers for choosing among existing RIASEC measures. In addition, we put forward important implications for developing future interest inventories with a basic interest structure.
Holland’s RIASEC Model
Holland’s RIASEC model provides a unified classification system for vocational interests and occupational environments (Campbell & Borgen, 1999; Campbell & Holland, 1972). Each of the RIASEC categories is a constellation of interests, preferred activities, beliefs, abilities, values, and characteristics, and can be used to describe both an individual’s interests and what their work environment supplies. Realistic (R) interests correspond to working with hands or tools (e.g., woodworking; mechanics). Investigative (I) interests are involved in scholarly or scientific activities and occupations (e.g., conduct physics research; psychologist). Artistic (A) interests correspond to being creative and expressive (e.g., paint a portrait; dancer). Social (S) interests are demonstrated in activities and occupations that involve helping or nurturing (e.g., nursing; elementary school teacher). Enterprising (E) interests entail business-related or influencing activities and occupations (e.g., open one’s own store; factory manager). Lastly, Conventional (C) interests are involved in work that is ordered and systematic (e.g., track shipping records; accountant).
Holland proposes that people tend to seek out work environments that supply interests based on their rank-ordered preferences for RIASEC types (Holland, 1985; 1997). As a career counseling tool, Holland’s RIASEC model is widely employed in career guidance, organizations (Su & Nye, 2017), and educational settings (Prediger & Swaney, 2004). The RIASEC model is also the dominant measurement framework for vocational interest research. For example, RIASEC interests predict people’s life goals (Stoll et al., 2020), influence academic and career choices (Hanna & Rounds, 2020; Usslepp et al., 2020), and predict job performance, satisfaction, and career success (Hoff et al., 2020; 2022; Nye et al., 2012; 2017; Rounds & Su, 2014; Van Iddekinge et al., 2011). Holland’s model also serves as the predominant framework for integrating interests with other individual difference domains, such as personality (Mount et al., 2005) and cognitive ability (Passler et al., 2015).
Assessing Convergence Among RIASEC Interest Inventories
With the widening application of Holland’s model, numerous measures have been developed to assess the RIASEC categories. However, a key question remains: to what extent do different interest inventories measure the same core content and provide people with similar results and career options? Or in other words, to what extent are RIASEC inventories interchangeable? Our study uses Holland’s theoretical model, and corresponding basic interests (Su et al., 2019), to investigate evidence for construct validity in three complementary ways. The first approach is based on correlations of RIASEC scores from different inventories using a multitrait-multimethod (MTMM) correlation matrix. Although other methods have been used to quantify convergent validity (e.g., regression coefficients, alerting correlation, and profile analysis), they are often built on correlation coefficients and therefore, would provide similar information to our first approach (Furr, & Bacharach, 2006). In addition, statistical models based on a MTMM correlation matrix can more effectively partition the separate influences of interest traits and inventory methods on RIASEC scores (Eid et al., 2006). The second approach uses high-point code convergence to examine a critical issue for practice: whether different inventories assign similar high-point codes (i.e., the most highly ranked interest scale) to individuals.
Together, these first two approaches offer valuable quantitative information on the degree of convergence among RIASEC measures. Yet neither of these two methods provides information about why interest scores do or do not converge from different measures. To understand the why, it is necessary to carefully examine the content coverage of each RIASEC scale within different measures. Hence, in a third approach, we adapt methods for establishing content validity and ask subject experts to examine item content similarities and differences using basic interests. We next discuss each approach in greater detail and propose hypotheses and research questions.
Correlation-Based Convergence
Cross-Correlation Between Similar or Same-Named Interest Scales.
Note. IP = Interest Profiler; SII = Strong Interest Inventory; ACT–VIP = ACT Vocational Interest Profile (1971–1974); SVIB = Strong Vocational Interest Bank; SCII = Strong-Campbell Interest Inventory; VPI = Vocational Preference Inventory; KGIS = Kuder General Interest Survey; OVIS = Ohio Vocational Interest Survey; CISS = Campbell Interest and Skills Survey; KOIS = Kuder Occupational Interest Survey; UNIACT-R = Revised Unisex ACT Interest Inventory (1995).
aMean within corresponding scales are calculated first before averaging across scales.
Only a few studies have assessed correlations among more than two interest inventories. Lowman et al. (2003) compared interest scores from a relatively diverse group of participants, including college students and job candidates, and they reported moderately high correlations (r = .61–.86, rmean = .73) among corresponding RIASEC scales measured by the Self-Directed Search, Vocational Preference Inventory, and Strong Vocational Interest Blank. In addition, Savickas et al. (2002) reported moderate correlations (r = .36–.72, rmedian = .59) among similar and same-named scales from five commonly used interest inventories that were developed by distinct research groups. Together, these two studies suggest fair convergent validity among interest measures, but unique variances still exist. However, these studies only focused on bivariate correlations and were not able to quantify the amount of shared variance among more than two measures. Thus, further research is needed to examine convergent validity among interest inventories using more advanced methods that provide a more holistic view.
In this study, we use the recommended set of Confirmatory Factor Analysis (CFA) models to assess convergent and discriminant validity evidence from multitrait-multimethod (MTMM) matrices (Campbell & Fiske, 1959). These models provide a method for estimating the independent effects (factor loadings) of method and trait latent factors on participant scores (Eid et al., 2006). Correlation-based convergence is assessed in these CFA models by comparing the variance explained by interest traits versus inventory methods. If interest traits explain considerably more variance in RIASEC scores compared to inventory methods, this provides evidence for convergence. Conversely, if inventory methods explain more variance in RIASEC scores than the traits, this suggests a lack of convergence (Widaman, 1985). Thus, with the first set of analyses, we test the following hypotheses:
Hypothesis 1a. In comparison to inventory method factors, interest trait factors will explain more variance in observed RIASEC scores.
Hypothesis 1b. Inventory method factors will explain a meaningful amount of unique variance in observed RIASEC scores above and beyond interest trait factors.
High-Point Code Convergence
High-point codes are a second important way of examining convergent validity among interest inventories. High-point codes, which reflect an individual’s strongest interest area, are widely used in both practice and research. In career guidance settings, high-point codes are often used to focus individuals’ career exploration efforts (Hanna & Rounds, 2020). High-point codes are readily interpretable for individuals across education levels, and they allow for direct linkage into occupations (Harmon et al., 1994). In research, high-point interests codes have been used to predict career choice (Hanna & Rounds, 2020) and outcomes within work environments, such as job satisfaction and performance (Hoff et al., 2020; Nye et al., 2017).
Only a few studies have examined the convergence of high-point codes from different inventories. We identified three studies that found that RIASEC high-point codes were not highly consistent across interest inventories (Harmon & Zytowski, 1980; Lowman et al., 2003; Savickas & Taber, 2006). In particular, Savickas and Taber (2006) showed that, between pairs of five interest inventories, exact three-letter high-point code matches only occurred at the frequency of 3–15%. Nevertheless, the authors reported concerns about the generalizability of these results since the participants primarily consisted of career counselors and researchers attending the Society for Vocational Psychology conference. Thus, it is important to revisit and further investigate high-point code convergence with a sample more representative of typical interest inventory users (Lowman, 2022).
Here, we examine the convergence of high-point codes by calculating the percentage of participants who obtain the same high-point code on two inventories (Rounds et al., 2021). High percentages of high-point code matches indicate that the inventories are likely to provide similar career exploration information to individuals. Although there is no widely accepted benchmark for the percentages that would support convergence, guidance is available from past research on matched high-point codes from different versions of the same inventory. Specifically, the hit rate between the 30- and 60-item O*NET Interest Profiler is 69.2%, and the hit rate between the 60- and 180-item versions is 78.4% (Rounds et al., 2021). Comparatively, we expect hit rates between different inventories to be lower. Therefore, we developed the following hypothesis:
In addition to the ceiling effect that can be inferred from past literature, we can also hypothesize about the lower bound of RIASEC high-point code convergence. Hanna and Rounds (2020) reported that on average, interest high-point codes have a 50.8% hit rate for predicting career choices. This meta-analytic hit rate denotes the criterion-related validity of RIASEC measures at large. In general, we expect the reliability of RIASEC high-point codes to exceed their criterion validity. Thus, we developed the following hypothesis:
Hypothesis 2: We expect that 50–70% of participants will receive the same RIASEC high-point codes from two different interest inventories.
Item Content-Based Convergence
A third way to assess validity is to systematically examine item content. Correlation-based convergent validity does not guarantee content validity (Dixon & Johnston, 2019), which means that reasonable score correlations between two measures do not necessarily indicate that they measure the intended underlying constructs. Each RIASEC category contains a wide range of work activities and environments that measure distinct basic interests (Su et al., 2019). If item content varies substantially across inventories, the underlying basic interest structure that comprises each RIASEC measure would also vary. When two same-name scales cover different sets of basic interests within the same construct domains, they could still be highly correlated at the between-person level, but the within-person RIASEC score rankings (i.e., interest profiles) could be sensitive to their content variability (Savickas & Taber, 2006). Since interest profiles are commonly used in career counseling settings, content variation could greatly affect practical indications derived from interest results. Despite the foundational importance of item content convergence for research and practice, content validity has received little attention in the field of interest measurement.
Item content is especially relevant in applied settings, such as career guidance, because the interpretation of an individual’s interest results depends on the specific content of an inventory. For example, the Realistic scale of one inventory could contain mainly construction and engineering items, while the Realistic scale from another inventory could include mainly agricultural and outdoor items. An individual who enjoys building or engineering but is not interested in raising dairy cows or working as a park ranger would score high on Realistic using the first inventory, but low on the second. Such differences in content coverage not only influence the client’s Realistic score but could result in a different high-point code leading to different career recommendations.
Basic interests offer a systematic and effective way to examine item content in interest inventories. Basic interests are more homogenous units of interest that can be organized into the structure of Holland’s RIASEC model (Rounds, 1995). Comparing the basic interest coverage of RIASEC scales provides a foundational understanding of what is commonly measured by different inventories, as well as their divergent coverage. If a basic interest is consistently measured by all four inventories, it is central to the construct space of its corresponding RIASEC scale. On the contrary, if RIASEC scales are composed of items from different basic interests, this would indicate a lack of consensus in their content coverage. In this study, we use a recently proposed basic interest framework—the Comprehensive Assessment of Basic Interests (CABIN; Su et al., 2019)—to investigate item coverage. We address the following question:
Research Question 1: To what extent do items in different interest inventories reflect the same basic interest structure within RIASEC scales?
Method
Data Collection
The current study incorporates two types of data: 1) participant survey responses from interest inventories and 2) expert ratings on content coverage of items from corresponding interest inventories. Participant survey responses were adapted from a prior data collection effort for [NAME BLINDED]’s dissertation, which involved the administration of multiple interest inventories to a sample of students that is more representative of typical interest inventory users. Our paper presents analyses based on that sample in addition to ([NAME BLINDED])’s efforts which added expert interest item ratings for the content analysis.
Interest Inventory Participants and Survey Administration
Participants were 327 undergraduates (approximately 80% were freshmen and sophomores) in a large midwestern university. All participants were enrolled in a career development course, an elective course for students who wanted to engage in career exploration and choice processes. Hence, these participants are likely to be representative of students who would typically complete an interest inventory. The course is not required and is open to students from all disciplines.
Participants ranged in age between 17 to 23 years (M = 18.86, SD = 1.18) and were relatively balanced in self-identified sex (43% male, 57% female). A large portion of the students were White/Caucasian (60.3%), followed by 18.4% Black/African American students, 5.3% Asian students, 4.1% multiracial students, .3% Native American students, and 5.6% of students reported “other” for race. In addition, 12.8% of students were Hispanic, and only one participant reported being an international student.
Students were given the set of paper-format surveys during a class meeting with written instructions to take home and complete at their convenience. They were given a date on which their completed packets would be picked up in class but not a particular order they needed to complete the measures. The researchers who took surveys to class for distribution also gave participants oral instructions, and participants were advised not to attempt to complete all the surveys in a single sitting, and to set the task aside if they experienced fatigue or found it difficult to remain focused.
Prior to data analysis, quality control items (embedded within the participants’ response forms and designed to detect random or inattentive responses) were checked. All 327 students in the career development course responded to the survey initially, but seven protocols were identified as suspects of careless responding, and were thus eliminated from the dataset. One participant did not complete the SDS, but other available responses were included in the dataset. This resulted in a final sample size of 319. Although our sample size is relatively small for detecting small correlations among latent factors in our proposed SEM models, post-hoc power analysis suggested that we have close to 1.00 power to detect factor loadings for all trait and method factors, which are the focus of our model result interpretation.
Measures
The present study focuses on four RIASEC measures: Self-Directed Search (SDS; Holland et al., 1994), ACT Revised Unisex Interest Inventory (UNIACT; ACT, 1995), Interest Profiler Short Form (IP Short Form; Rounds et al., 2010) and Strong Interest Inventory (SII; Harmon et al., 1994). These four inventories were chosen because they are among the most used interest inventories, and because they all include a set of scales that specifically measure the RIASEC constructs. Study participants granted permission for their interest inventory results from class to be combined with their responses to the SII, which was administered outside of class. Out of these four measures, we were able to obtain items from the SDS, UNIACT, and IP, but not the SII which is a proprietary measure that does not release its items.
Self-Directed Search (SDS; Holland, 1985)
The SDS was designed by Holland to be a cheap, practical vocational guidance tool that allows users to self-administer, self-score, and self-interpret (Holland et al., 2020). In the SDS, each RIASEC construct is evaluated by 11 activity items in Like/Dislike response format (e.g., repair cars), 11 Competencies items in Yes/No response format (e.g., I can repair furniture), 14 Occupations items in Yes/No response format (e.g., social worker), and 12 Self-Estimates items scoring from 1 (low) to 7 (high) (e.g., scientific ability). An advantage of the dichotomous scoring system is that users can simply add up the “Like” and “Yes” responses for each RIASEC to obtain their interest scores, which requires no transformation or norming on the scores. Hence, participants’ raw scores were also used for their SDS results in this study, and these scores were used to identify their high-point interest code. Scale reliabilities are presented in the inventory manual, and alphas for RIASEC constructs ranged from .83 to .92 (Holland, 1985).
UNIACT-R Level 2 (UNIACT; Swaney et al., 1995)
The UNIACT Level 2 was designed to measure interests for college students and adults, and to only include sex-balanced items. Although the UNIACT effectively reduces sex differences in interest scores, it can reflect a narrower range of RIASEC domains that measure interest in activities with small sex differences (Su et al., 2009). The UNIACT-R contains 90 activity items in total, with 15 items for each RIASEC scale (e.g., operate office machines). All items use the Dislike/Indifferent/Like response format. The UNIACT uses slightly different names for the RIASEC constructs than Holland’s, but the scales are reportedly equivalent: R = Technical, I = Science, A = Arts, S = Social Service, E = Business Contact, C = Business Operations. The UNIACT was scored following instructions in the UNIACT Technical Manual (Swaney et al., 1995). Raw score values range from 15 to 45 and are converted to standardized T-scores (M = 50, SD = 10). Reliability of scales is reported in the manual, and alphas ranged from .77 to .91 (Swaney et al., 1995).
Interest Profiler Short Form (IP Short Form; Rounds et al., 2010)
The IP was developed as one of the Career Exploration Tool sponsored by the U.S. Department of Labor (DOL). One major advantage of the IP is that it is publicly accessible, and the results are directly connected to O*NET’s occupation data for automatically providing users with career recommendations. The IP Short Form contains 60 activity items in total, with 10 items for each RIASEC construct (e.g., manage a retail store). All items use the Dislike/?/Like response format. Raw score values range from 0 to 120. We note that during [NAME BLINDED]’s original data collection, all 180 Interest Profiler items were administered (Lewis & Rivkin, 1999). However, the IP Short Form is the primary form employed in current practice, so we derived the corresponding scores for the Interest Profiler Short-Form using the instructions found on O*NET (Rounds et al., 2021). The reported scale reliability information for the IP Short Form ranged from .78 to .87 (Rounds et al., 2010).
Strong Interest Inventory, (SII; Harmon et al., 1994)
Different from other inventories included in this investigation, the SII was not originally developed to measure general interests like the RIASEC. It started with empirical occupational interest scales which compare the responses of an individual’s interests to those of satisfied employees in specific occupation groups. To make interest results easier to interpret and generalize, basic interest scales were added to the Strong in 1969 (Campbell, 1971). Later, the 1994 edition SII further incorporated Holland’s RIASEC types as an organizing system for its basic and occupational interests. Although the SII has the longest history and has made a profound impact on interest measurement, it can be costly for many job seekers. In addition, the SII has a protected item pool which limits its accessibility to researchers.
The SII evaluates interests using eight types of items and 317 items in total. Five item types are in the Dislike/Indifferent/Like response format: Occupations (135 items), School Subjects (39 items), Activities (46 items), Leisure Activities (29 items), and Types of People (20 items). Two item types ask participants to choose preferences between several options, including Activities (30 items) and Preferences in the Work World (6 items). Personal Characteristics items (12 items) are in the Yes/?/No response format. SII scale score data were taken from reports scored by the publisher, where raw score values were converted to standardized T-scores (M = 50, SD = 10; Harmon et al., 1994). Reliability estimates for RIASEC scales are reported to range from .90 to .94 (Harmon et al., 1994).
Data Analysis
Correlation-Based Convergent Validity
To test hypotheses 1a and 1b, we conducted a nested model comparison among four sets of confirmatory factor analysis (CFA) models on the multitrait-multimethod (MTMM) matrix. Prior to the analyses, we treated random item-level missingness by using full information maximum likelihood to estimate a scale-level correlation matrix with the lavaan package in R. This correlation matrix was then used as data input for further modeling estimation. Model fit for all estimated models was evaluated using comparative fit index (CFI), Tucker–Lewis index (TLI), root-mean-square error of approximation (RMSEA), and standardized root-mean-square residual (SRMR). CFI and TLI values greater than .95, and RMSEA and SRMR values less than .05, suggest good/close model fit; CFI and TLI values greater than .90, RMSEA and SRMR values between .05 to .08 suggest fair/reasonable model fit; and models with RMSEA .08 and .10 indicate mediocre fit (Browne & Cudeck, 1992; Hu & Bentler, 1999).
To compare models, we first fitted a correlated-methods (CM) and a correlated-traits (CT) model. The CM model assumes that variances in observed interest scores are influenced by common latent inventory method factors and unexplained residual variances, whereas the CT model is built on the same assumption regarding latent interest trait factors. Variances across RIASEC scores in the same inventory are assumed to load on one common method factor in the CM model; variances in Realistic scores from all four inventories are assumed to load on one common Realistic trait factor in the CT model. These two models are estimated to compare with fit statistics of models that estimate both trait and method effects. Next, we fit a correlated traits-correlated methods model (CT-CM; Marsh & Grayson, 1995; Widaman, 1985), which assumes that variances in observed interest scores are influenced by both an interest trait and the inventory to which they belong. Traits were allowed to freely correlate, as were methods. No correlations are allowed between trait and method factors to estimate the independent effect of each source. Figure 1 shows a graphical representation of the CT-CM model. Model fit statistics are compared between the CT and CT-CM model to indicate the importance of method effects. Correlated Trait – Correlated Method (CTCM) Model of RIASEC Inventory Convergence. Note. SDS = Self-Directed Search; IP = Interest Profiler (Short Form); UNIACT = the Unisex ACT Interest Inventory; SII = Strong Interest Inventory. M1 to M4: different inventories 1 to 4.
Although the CT-CM model separately estimates trait and method effect, the influence of a general trait factor can be confounded by correlated method factors (Eid et al., 2008). Hence, we estimated the third set of CFA models—correlated traits-correlated methods-1 models (CT-C (M-1); Eid et al., 2003)—that offer an alternative view on interest trait and inventory method effects. Similar to the CT-CM model, these models estimate both trait and method effects. Unlike the CT-CM model, the CTC(M-1) model does not suffer from the existence of a potential general factor that can inflate the explained variance of method factors (Eid et al., 2006). The CT-C (M-1) model requires choosing one referent method as the standard for comparison. Because we do not assume one interest inventory to be the gold standard, we estimated four CT-C (M-1) models that each treat one inventory as the referent method. All models were specified and estimated using Mplus version 8.1 (Muthén & Muthén, 1998-2017).
High-Point Code Based Convergence
To test hypotheses 2a and 2b, we examined the level of agreement among high-point codes for each participant on each measure. High-point codes were assigned according to instructions in each inventory’s manual. In general, this involved assigning a letter code based on the person’s highest scores. Tied scores were randomly split to provide the respondent with a single high-point code. In addition to percentages of matched high-point codes between each pair of inventories, we also assessed agreement rates using Cohen’s Kappa (K; Cohen, 1960) and Fliess’s Kappa to better understand the degree of high-point code convergence. These two statistics take into account the chance probability of high-point code agreements (Fleiss, 1981). Kappa values above .60 indicate substantial agreement, values between .41 and .59 indicate fair agreement, and values between .21–.40 indicate poor agreement (Fleiss, 1981).
Item Content-Based Convergent Validity
To address Research Question 1, we first identified content coverage within each RIASEC measure. Traditionally, content validity is established at the beginning of the scale development process by asking subject matter experts to sort items into potential constructs (Hinkin & Tracey, 1999). We adapted this approach to assess convergent validity in content coverage across measures, and four expert raters (two I/O psychology professors and two graduate students) were asked to individually categorize all 366 items from three interest inventories (SDS, UNIACT, IP) into basic interest dimensions. The Strong Interest Inventory items are not open to the public domain and therefore were not included in content analyses.
For the most thorough and updated representation of the interest space, we used the Comprehensive Assessment of Basic Interests (CABIN; Su et al., 2019) which includes 4 to 10 basic interests for each RIASEC domain. The Realistic domain covers 10 basic interests (Agriculture, Animal Service, Athletics, Construction, Engineering, Mechanics/Electronics, Outdoors, Physical/Manual Labor, Protective Service, Transportation/Machine Operation), Investigative covers 4 (Life Science, Mathematics/Statistics, Medical Science, Physical Science), Artistic covers 7 (Applied Arts & Design, Creative Writing, Culinary Art, Media, Music, Performing Arts, Visual Arts), Social covers 8 (Health Care Service, Human Resources, Humanities & Foreign Language, Personal/Service, Religious Activities, Social Science, Social Service, Teaching Education), Enterprising covers 8 (Business Initiatives, Law, Management/Administration, Marketing/Advertising, Politics, Professional Advising, Public Speaking, Sales), and Conventional covers four basic interests (Accounting, Finance, Information Technology, Office Work).
After each rater coded the items independently, acceptable (k > .60) to good (k > .80; Gelfand & Hartmann, 1975; Landis & Koch, 1977) inter-rater agreement was found on all interest types, ranging from k = .75 for Conventional items to k = .96 for Investigative items. Raters met after finishing their coding and discussed items that received different categorizations. The final categorization for each item was determined based on the majority vote. A few items had equal votes for two separate categories, and they were counted as half an item to each category.
Results
Descriptive Statistics on Interest Scale Scores.
Note. SDS = Self-Directed Search; IP = Interest Profiler (Short Form); UNIACT = the Unisex ACT Interest Inventory; SII = Strong Interest Inventory. R = Realistic; I = Investigative; A = Artistic; S = Social; E = Enterprising; C = Conventional.
Multitrait-Multimethod (MTMM) Matrix.
Note. IP = Interest Profiler; SD = Self-Directed Search; U = Revised Unisex ACT Interest Inventory; SI = Strong Interest Inventory.
How Much Do Trait and Method Factors Affect RIASEC Scale Scores?
Fit Statistics for Competing Confirmatory Factor Analysis Models.
Note. SDS = Self-Directed Search; IP = Interest Profiler (Short-Form); UNIACT = the Unisex ACT Interest Inventory; SII = Strong Interest Inventory; CT = Correlated Trait model; CTCM = Correlated trait-Correlated method model; CTC(M-1) = Correlated trait-Correlated method minus one model.
To test Hypothesis 1b that method factors would explain a meaningful amount of variance in RIASEC scores, a nested model comparison was conducted. The CTCM model (RMSEA = .09, CFI = .93, TLI = .90, SRMR = .06) performed significantly better than trait-only CT model (Δχ2 (30, N = 219) = 841.71, p < 0.01). This suggests that method factors explain a substantive amount of unique variance in observed RIASEC scores, supporting Hypothesis 1b. In addition to its better model fit, the CTCM model also does not force interpretations to be based on only one interest inventory as the gold standard, but rather how well each inventory assesses each interest type. For these reasons, we retain and discuss factor loadings of the CTCM model.
Correlations among Latent Trait and Method Factors.
Note. SDS = Self-Directed Search; IP = Interest Profiler (Short Form); UNIACT = the Unisex ACT Interest Inventory; SII = Strong Interest Inventory. R = Realistic; I = Investigative; A = Artistic; S = Social; E = Enterprising; C = Conventional.
Standardized Parameters for the Correlated Traits-Correlated Methods (CTCM) Model.
Note. SDS = Self-Directed Search; IP = Interest Profiler (Short Form); UNIACT = the Unisex ACT Interest Inventory; SII = Strong Interest Inventory. R = Realistic; I = Investigative; A = Artistic; S = Social; E = Enterprising; C = Conventional.
Do Respondents Receive the Same High-Point Codes Across Interest Inventories?
Percentage of High-point Code Match and Cohen’s k Agreement Indices Among Inventories.
Note. Fleiss’ kappa for overall agreement k = .39. Above diagonal are the percentage of participants who have matched high-point codes between pairs of inventories. Below the diagonal are Cohen’s k index of agreement. k values between .41–.59 indicates fair agreement whereas k values between .21–.40 indicates poor agreement.
Correspondingly, there was poor (k = .34) to moderate (k = .45) Kappa agreement rates between pairs of inventories (M = .40). Only two pairs—SDS-SII (63%; k = .51) and SDS-IP (61%; k = .49)—showed fair agreement. The UNIACT-SII pair showed the lowest level of agreement (40%; k =.27). We calculated Fleiss’ kappa (Fleiss, 1981) for multiple-rater (inventory) agreement, which also indicated poor high-point code agreement between the four inventories (k = .39). Thus, the high-point code comparison revealed weaker evidence for convergent validity compared to the correlational analyses. We next conducted item content analyses for the SDS, IP, and UNIACT to investigate whether the RIASEC scales from different inventories contain different structures of basic interests.
Do RIASEC Inventories Share Similar Item Content Coverage?
Item Count for Basic Interest Categorization and Interrater Agreement.
Note. SDS = Self-Directed Search; IP = Interest Profiler (Short Form); UNIACT = the Unisex ACT Interest Inventory.

Averaged Percentage of Items Belonging to Basic Interests. Note. This figure visualizes the average row percentage of basic interest categorization that is reported in Table 8. Sci. = Science; Serv. = Service; Prof. Advise = Professional Advise; Pub. Speak = Public Speaking; IT = Information Technology.
The rest of the construct space consists of peripheral basic interests distributed unevenly across the three inventories. For example, for measuring Investigative, the IP places more emphasis on Medical Science, and the UNIACT places more emphasis on Life Science. The SDS and UNIACT also contain more broad items that could not be categorized into basic interests, especially for Investigative (e.g., read scientific books or magazines), Social (e.g., I find it easy to talk with all kinds of people), and Enterprising scales (e.g., influence others). These items were put into the “other” category for each scale.
Summary of Content Analysis Findings.
Note. Core Basic Interests are measured by 10% or more of items in all three inventories (SDS, UNIACT, IP); Peripheral Basic Interests are measured by 10% or more of items in two of the three inventories; Under-Covered Basic Interests are measured by 10% or less of items in at least two inventories, and if a basic interest is measured by 10% or more items in one inventory, the specific inventory is noted in parentheses.
In terms of under-covered basic interests, 17 out of the 41 basic interest dimensions were missing or insufficiently covered in at least two of the inventories. For Realistic, there was no coverage of six basic interest scales: Engineering, Agriculture, Physical/Manual Labor, Athletics, Protective Services and Animal Services. For Investigative, only the SDS assessed Mathematics/Statistics. For Artistic, Culinary Art was not measured. For Social, Human Resources, Humanities & Foreign Language, and Religious Activities were missing, while only the SDS contained multiple items measuring Social Science and only the UNIACT covers Personal Service. For Enterprising, only the UNIACT measured Professional Advising, whereas only the SDS measured Public Speaking. For Conventional, Information Technology was missing in all but the IP. Overall, these content gaps highlight new basic interest scales which could be included in future interest inventories. The results also reveal both similarities and differences in the basic interest coverage across RIASEC scales, with important implications for understanding what is measured by existing interest inventories.
Discussion
Vocational interest research and practice rest on a foundational assumption that different inventories assess the same latent constructs and provide similar career exploration information to individuals. The current study critically examined this assumption using four major interest inventories: Self-Directed Search (SDS; Holland et al., 1994), ACT Revised Unisex Interest Inventory (UNIACT; ACT, 1995), Interest Profiler Short Form (IP Short Form; Rounds et al., 2010), and Strong Interest Inventory (SII; Harmon et al., 1994). We used three distinct methods to assess different aspects of construct validity: 1) confirmatory factor analysis (CFA) modeling of multitrait-multimethod (MTMM) matrices, 2) high-point code agreements, and 3) item content coverage. Results from the first two analyses revealed that RIASEC scores from different measures were generally highly correlated, but there was lower than expected high-point code agreement between pairs of inventories. Item content analysis provided one possible explanation for the lack of high-point code convergence by showing that despite interest inventories covering some core basic interests, there are important differences in the content of distinct inventories. We next discuss these findings in greater detail and present theoretical and applied implications.
RIASEC Scores Converged, but High-Point Codes Often Diverged
The first two sets of analyses each painted a distinct picture of convergent validity evidence. First, we estimated correlated traits-correlated measures CFA models, which revealed that interest trait factors had stronger influences on RIASEC scores compared to inventory method factors. This provides evidence for convergence among the SDS, IP, UNIACT, and the SII. Nonetheless, the influence of methods factors on interest scores was also substantial—models with both trait and method factors had a substantially better fit than the trait-only model. The rejection of the trait-only model suggests the final observed score is dependent not only on the trait assessed but also on the inventory used to assess the trait (Eid et al., 2006).
The second method, high-point code analysis, revealed weaker evidence for convergence. High-point codes, reflecting a person’s strongest RIASEC interest, are widely used in vocational interest research and practice (Hanna & Rounds, 2020; Hansen, 2019). Our results indicated that if an individual took two RIASEC inventories, they would be assigned distinct high-point codes about half the time (Magreement = 52%). This degree of convergence is comparable to the average criterion validity (hit rate) of RIASEC inventories in predicting career choices (Hanna & Rounds, 2020), but is substantially lower than that observed from interest measures using different subsets of the same item pool (69–78%; Rounds et al., 2021).
Item Content Analyses Revealed Convergence, Discrepancies, and Missing Domains
Item content is the basic unit that defines any latent psychological construct. Yet, it has received little attention in the interest assessment literature. Our content analysis distinguished RIASEC scales into three parts: core, peripheral, and missing basic interests. The IP, the SDS, and the UNIACT mostly agreed on the core basic interests within each RIASEC scale. This suggests that interest inventories have a broad consensus on what work activities and occupations are most typical of each RIASEC scale. As noted, Figure 2 provides an overview of common basic interests within each RIASEC scale. This overview serves as a useful tool for researchers to understand what is currently being measured by major RIASEC inventories.
Apart from the core basic interests, the three inventories also included distinctly different basic interest dimensions. We identified these scales as peripheral basic interests. Discrepancies in peripheral basic interests can lead to different overall representations of the RIASEC constructs, potentially explaining the low high-point code convergence among these inventories. For example, respondents who enjoy physical science and mathematics, but are less interested in life or medical science, may receive a higher investigative score using the SDS compared to UNIACT. These differences in scores can lead to variability in ranking RIASEC scores, especially when a person shares similar levels of interest in multiple types. In other words, content differences can add up and ultimately give respondents different interest results.
We also found that a large number of basic interests (17 of 41) are completely missing or insufficiently represented in current RIASEC measures. One possible reason for the incomplete representation of basic interests is that some basic interests reflect more than one RIASEC category. For example, Social Science captures both Investigative and Social types (Su et al., 2019); and Human Resources captures a mix of Social and Enterprising. Including basic interests that tap into more than one RIASEC category can affect the structure of Holland’s model, which is an important quality in RIASEC inventories (Rounds & Day, 1999). However, as a growing number of jobs require an intersection of multiple interest types, ignoring basic interests hinders the effectiveness of interest inventories for providing career guidance.
Implications for Administrating RIASEC Inventories and Measuring Basic Interests
Our content analysis offers applied implications for interpreting results from existing RIASEC inventories and constructing future inventories using basic interests as a structural unit. RIASEC inventories are not interchangeable, and practitioners can benefit by considering the characteristics of target clients for choosing a RIASEC inventory to administer, especially if individuals expressed high interest in areas outside of the core basic interest dimensions (Lowman, 2022). The divergent results from our two quantitative analyses suggest that interest inventory users should be cautious when interpreting high-point codes. Although many practitioners or inventory instructions indeed walk clients through their entire interest profiles, it is also important to communicate the complexity of individual interests and the fact that people rarely fit neatly into one interest type (Lowman, 2022). Hence, interest inventory users should be reminded that interest assessments are not diagnostic, and they need to remain flexible in thinking about their interests. Flexible thinking is particularly important when users take multiple self-interpreted interest measures—they need to be made aware that if they receive different interest codes from two inventories over time, it does not necessarily reflect true changes in their interests. To help individuals correctly interpret and further understand their interest results, existing RIASEC assessments should include relevant information about what basic interests are being measured and those not being measured under each construct.
It is also important for researchers to consider differences between RIASEC inventories when including interest scales in studies. Differences in item content from inventories can yield different scale scores and consequently different study results. For example, previous research suggests variability between interest inventories when examining their fit to Holland’s circumplex, which results in different degrees of structural validity (Armstrong et al., 2003; Rounds & Tracey, 1996; Tracey & Rounds, 1993). Previous meta-analyses on the criterion-related validity of vocational interests also suggest differences in validities obtained from different measures. For example, the relations between interest fit with task performance (ρ = .20 to .31; Nye et al., 2017) and interest fit and job satisfaction (ρ = .11 to .24; Hoff et al., 2020) vary depending on the inventory. Some variation of the observed differences in criterion validity of interest inventories can be explained by each measure’s unique history. For example, the SII was developed to emphasize predictive validity using empirically derived occupational scales. However, variation in item content coverage can impact all aspects of validity for an inventory. Hence, researchers should be aware of what basic interests are included by different inventories when selecting the most relevant measure. For example, using only the IP to assess the interests of financial analysts would fail to assess the basic interests most likely to be relevant to the sample (i.e., finance and statistics/mathematics). To address this gap, the researcher could supplement the interest measure with finance and statistics basic interest scales from a public domain basic interest inventory (e.g., Liao et al., 2008).
For both practitioners and researchers who plan to assess basic interests, we offer the following recommendations for selecting basic interest scales according to the targeted respondents. To provide career or educational guidance for students, we suggest including core basic interests with the addition of basic interests that reflect important instructional programs. For example, including all basic interests under the Investigative type is desirable because each can be directly linked to college majors. To offer career guidance for the general public, we suggest including basic interests that map onto fast-growing occupations in the workforce. For instance, according to the U.S. Bureau of Labor Statistics (BLS), retail trade and healthcare each makeup roughly 10% and 12% of the service-providing industry (U.S. Bureau of Labor Statistics, 2021). Therefore, it is important to include sales and healthcare service dimensions in Enterprising and Social scales, respectively. Another example is Information Technology, which connects to conventional interests and is projected to grow 12% by 2029 (U.S. Bureau of Labor Statistics, 2022).
Altogether, our results emphasize the importance of assessing convergent validity evidence in multiple, distinct ways. CFA modeling on the MTMM matrix effectively compares the independent effects of traits and methods on scores, and they help assess whether different inventories are measuring the same latent constructs (Eid et al., 2008). Yet, they provide little information about whether different inventories yield the same results in practice. Inventories may be fairly robust in measuring the same underlying constructs, but subtle variations in item content can yield divergent results that lead to different high-point codes. Construct validity research should therefore take into consideration how scores are used and interpreted in practice. Otherwise, the real-world impact of using different measures to assess the same construct could be masked by high correlation-suggested convergence.
Strengths, Limitations, and Future Directions
To our knowledge, the current study is the most comprehensive assessment of convergence and content validity among vocational interest inventories. We quantified the degree of convergence among major interest inventories using correlation-based and high-point code-based approaches, and we also examined item content similarities and differences across RIASEC scales. Each method revealed distinct aspects of convergent validity that provide important implications for understanding what is being measured by RIASEC inventories. However, there are several noteworthy limitations of the current research.
First, for our quantitative analyses, we adapted prior data on four major interest inventories. Although our adaptation allows us to include the most widely used public domain inventory today—the Interest Profiler Short Form, there are other interest inventories that we did not analyze (Hansen, 2019). In addition, we were unable to obtain item-level data from the Strong Interest Inventory and calculate its scale reliabilities in our study. Future researchers who have access to all items from multiple interest measures could replicate our quantitative analyses on the item-level data. Furthermore, future content validity studies on interest measurement could benefit from analyzing a wider selection of publicly accessible interest items, particularly in examining the consistency of core basic interest dimensions identified in the current study.
Second, in addition to differential coverage, convergence also depends on other factors relevant to the development of each inventory. For example, interest inventories can differ in their item selection to reduce sex differences, the use of normed versus raw score reporting, items written for different levels of education, and item type (e.g., occupations, activities, and school subjects). Notably, the UNIACT selected items with small sex differences, which potentially explains its stronger method effect on Realistic and Social, the two scales with the largest sex differences (Su et al., 2009). Future studies could further assess the effects of specific elements on convergence among interest scores.
Third, there is currently no uniform consensus on the number and structure of basic interests. Hence, our content analysis results, which are based on the Comprehensive Assessment of Basic Interest (CABIN; Su et al., 2019), should be interpreted as constructive, not definitive. With these limitations in mind, we put forth our key message for improving interest assessment: As the world of work changes, vocational interest inventories must update their representation of interest dimensions through bottom-up processes, capturing the types of jobs most relevant in the present and future labor market. In addition, all methods of construct validity, including content and convergent, should be regularly assessed to keep interest inventories up to date. Thus, this study provides a benchmark for future revisions on the construct validity of interest inventories.
Conclusion
The current study assessed construct validity among major interest inventories using three methods, each providing unique contributions to better understand interest measurement. The results showed that major interest inventories reflect similar RIASEC traits in general, but they less often provide consistent high-point codes to respondents. Moreover, our analysis explored why interest inventories are not interchangeable by showing their distinct coverage of peripheral basic interests and missing basic interests. Overall, our findings help integrate the theoretical understanding of Holland’s (1997) RIASEC constructs with important implications for improving existing RIASEC inventories, as well as developing the next generation of basic interest-based assessments.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Author Note
Chu Chu and Marry T. Russell share joint first authorship of this manuscript. Our study utilizes Mary Russell’s dissertation (Russell, 2008), Assessing Vocational Interests: Convergence and Divergence of Inventories and Informants, as a starting point. Chu Chu conducted reanalyses of Russell’s dataset, led efforts to expand the scope of the paper to include item content analysis of basic interests, and prepared this manuscript for publication.
