Abstract
The validity and utility of translated instruments (psychological measures) depend on the quality of their translation, and differences in key linguistic characteristics could introduce bias. Likewise, linguistic differences between instruments designed to measure analogous constructs might contribute to similar instruments possessing dissimilar psychometrics. This article introduces and demonstrates the use of natural language processing (NLP), a subfield of artificial intelligence, to linguistically analyze 13 translations of two psychological measures previously translated into numerous languages. NLP was used to generate estimates reflecting specific linguistic characteristics of test items (emotional tone/intensity, sentiment, valence, arousal, and dominance), which were then compared across translations at both the test- and item-level, as well as between the two instruments. Results revealed that key linguistic characteristics can profoundly vary both within and between tests. Following a discussion of results, the current limitations of this approach are summarized and strategies for advancing this methodology are proposed.
The vast majority of self-report measurement instruments (i.e., tests) are constructed using a single language, requiring instrument translation before these can be used in research and practice with populations who communicate in other languages. The psychometric properties of these instruments across languages depend on accurate and culturally applicable translations, with translation quality conventionally demonstrated with traditional psychometric studies (e.g., criterion-related validity, measurement invariance). However, advances in the accessibility and validity of psychological assessment across languages and cultures require an increasingly better understanding of the similarities and differences between instrument translations because characteristics of test item wording, format, and context influence the way individuals respond on self-reports (Schwarz, 1999). One potentially important feature to examine is the degree of variability in key linguistic characteristics across translations. Most importantly, attention should be given to semantic and pragmatic differences between translations of the same measure that might influence how an individual responds to test items. Yet, the importance of understanding semantic and pragmatic differences between tests likely extends beyond the context of instrument translation, as these linguistic differences might also contribute to the divergent psychometric properties observed between instruments designed to measure the same latent construct.
Linguistic differences between instruments and test translations can feasibly introduce unintended measurement bias through multiple mechanisms. First, semantics and pragmatics are important components of language for the accurate communication of a fact or sentiment (Crystal, 1981). Consider the statements, “I’m sad” and “I’m depressed,” both of which generally communicate a negative emotional experience. However, the stronger negative sentiment transmitted by the second sentence—in this case, through the use of the word “depressed” instead of “sad”—changes the intensity of this statement and, consequently, changes what it would mean for a test responder to endorse this item. Most psychologists and psychometricians would agree that replacing the word “sad” with the word “depressed” is problematic because these words are not equivalent, and their exchange would change the latent meaning and intensity of the test item. However, determining the equivalence of words and multi-word items across languages is not as simple as the offered example (see Regmi et al., 2010). As such, translated items that are presumably equivalent might not convey the same meaning across languages and cultures.
Priming effects are a second mechanism through which linguistic differences might introduce measurement bias. Priming occurs when stimuli function as subtle contextual cues (primes) that increase the accessibility of mental concepts associated with the presented cues (Bargh & Chartrand, 2000). Studies repeatedly demonstrate the effects of lexical and semantic priming (Meyer & Schvaneveldt, 1971) on participants’ responses to psychological measures (e.g., Bartz & Lydon, 2004; Stapel & Blanton, 2004; Stapel & Koomen, 2000, 2001). Pertinent for this article, the sentiment and emotional tone/intensity of an item conceivably functions as a subtle contextual cue that influences the way an individual responds to the test item, and possibly successive items. These priming stimuli, which operate through multiple channels (e.g., cognition, perception, affect, motivation; Elgendi et al., 2018; Molden, 2014), can generate assimilating or contrasting effects: A test response can be defined as the product of an assimilating effect when the response distortion aligns with the priming stimulus (e.g., a test item with a sad emotional tone results in an individual rating themselves as more depressed than would be the case in the absence of the priming effect), whereas a contrasting effect induces a response distortion that diverges from the priming stimulus (e.g., when an individual denies ordinary experiences of sadness to distance themselves from an item asking about depression that has an overall negative sentiment).
Regardless of the mechanism(s) through which linguistic differences might be introducing measurement bias, the fact of the matter is that certain linguistic differences might, at least partially, be responsible for noninvariance seen across translations of psychological measures (e.g., Le Corff et al., 2022; Natoli et al., 2022; Sorrel et al., 2021). As noted above, the importance of understanding linguistic differences between tests might also underlie the divergent psychometrics observed when comparing tests that are expected to be psychometrically similar for other reasons (e.g., tests intended to measure the same latent construct) (Podsakoff et al., 2003). Before this potential issue can be addressed, it’s first necessary to understand the nature of variability in linguistic characteristics across instrument translations and between measures assessing analogous constructs and optimize a methodology appropriate for these lines of inquiry. Natural language processing (NLP), a subfield of artificial intelligence dedicated to helping machines learn to interpret, manipulate, and understand human language, offers a promising technology for meeting this objective.
Natural Language Processing
With psychological tests predominantly operating through written and spoken language, NLP is a uniquely capable technology for analyzing and comparing instrument translations. The utility of NLP can be categorized into two areas: natural language understanding and natural language generation, the first of which is more applicable to the current study. Paralleling the field of linguistics, natural language understanding analyzes the phonology, morphology, syntax, semantics, and pragmatics of text data, enabling machines to understand and analyze (parts of) language. The present study focuses on understanding the semantic and pragmatic levels of language, using NLP to recognize the “meaning” of a given test item and quantify its sentiment (i.e., the positivity or negativity expressed) and emotional tone/intensity. This is done by deconstructing each test item into its parts (i.e., words, called tokens), assigning each token a value representing the sentiment or emotion(s) represented and their intensity, and then aggregating these at the sentence (item) level. The resulting output delivers a vector of values for each test item representing sentiment, the presence and degree of different emotions, and their corresponding valence defined by validated lexicons. These results can then be analyzed and compared across items, tests, and translations.
The Current Study
This article introduces and demonstrates the use of NLP to linguistically analyze two measures of personality dysfunction that have been translated into numerous languages and validated: the Level of Personality Functioning Scale—Brief Form 2.0 (LPFS-BF 2.0; Weekers et al., 2019) and the Personality Disorder Severity ICD-11 Scale (PDS-ICD-11; Bach et al., 2021). To do so, test items from 13 translations of these instruments were subjected to emotion recognition and sentiment analysis using multilingual NRC Word-Emotion Association (NRC WEA; Mohammad & Turney, 2013), NRC Emotion Intensity (NRC EI; Mohammad, 2018a), and NRC Valence, Arousal, and Dominance (NRC VAD; Mohammad, 2018b) lexicons. This application of NLP was used to generate values reflecting specific linguistic characteristics of each test item (i.e., emotional tone/intensity, sentiment, valence, arousal, and dominance), which were then compared across translations at both the test- and item-level. As a secondary objective, these same linguistic characteristics were compared across the two test instruments to quantify (dis)similarity in the emotional tone/intensity, sentiment, valence, arousal, and dominance of different measures intended to assess equivalent latent constructs. As a whole, this study aimed to introduce and demonstrate a methodological approach for assessing variability in linguistic characteristics across instruments and across instrument translations.
Method
Test Instruments
Level of Personality Functioning Scale—Brief Form 2.0
The LPFS-BF 2.0 (Weekers et al., 2019) is a 12-item self-report measure of self-functioning and interpersonal functioning (i.e., personality functioning) impairment. Items 1 to 6 reflect aspects of self-functioning and items 7 to 12 reflect aspects of interpersonal functioning. Individual self-functioning impairment and interpersonal functioning impairment scores are calculated by aggregating the six items comprising each scale, and a total personality functioning impairment score is calculated by summing all 12 items. The LPFS-BF 2.0 has demonstrated strong psychometric properties when administered in several languages, including Danish, Dutch, English, Estonian, French, German, Italian, and Spanish (Combaluzier et al., 2023; Cottin et al., in preparation; Gritti et al., in preparation; Le Corff et al., 2022; Natoli et al., 2022; Oitsalu et al., 2022; Spitzer et al., 2021; Weekers et al., 2019, 2023; Zimmermann et al., 2020).
Personality Disorder Severity ICD-11 Scale
The PDS-ICD-11 (Bach et al., 2021) is a 14-item self-report measure designed to assess global personality disorder severity as operationalized in the chapter on personality disorders and related traits in ICD-11. Test items reflect the literal concepts and wordings of the ICD-11’s framework for personality disorder severity determination and can be summed to produce an overall severity index. The PDS-ICD-11 has demonstrated strong psychometric properties when administered in multiple different languages, including Danish, English, and Spanish (Bach et al., 2021; Gutiérrez et al., 2023; Natoli & Rodriguez, 2024).
At the time of data analysis, the LPFS-BF 2.0 and PDS-ICD-11 had both been translated into the following languages that use Latin character sets 1 : Czech, Danish, Dutch, English, Estonian, French, German, Italian, Norwegian, Polish, Portuguese, Spanish, and Turkish.
Lexicons
Three existing manually curated multilingual lexicons that are publicly available were used for the current study. The NRC EI (Mohammad, 2018a) lexicon consists of over 10,000 terms associated with emotions to varying degrees. Each term in the lexicon has eight corresponding intensity scores for eight basic emotions (anger, anticipation, disgust, fear, joy, sadness, surprise, and trust), which vary from 0 (conveys the lowest amount of a given emotion) to 1 (conveys the highest amount of a given emotion). For instance, the term “disagreement” would be assigned an anger intensity score of .348, a sadness intensity score of .438, and an intensity score of .000 for the remaining six emotions. The lexicon’s intensity scores were derived by manual annotators following a standardized annotation scheme and best–worst scaling procedure (see Mohammad, 2018a). The NRC WEA (Mohammad & Turney, 2013), containing approximately 14,000 terms, and the NRC VAD (Mohammad, 2018b), containing over 20,000 terms, were created following similar procedures as the NRC EI. These two lexicons were used in the current study to generate positive/negative sentiment values and valence, arousal, and dominance values, respectively. Positive and negative sentiment is assigned to each term in a test item dichotomously (1 = present, 0 = not present) whereas assigned values of valence (positive—negative), arousal (active—passive), and dominance (full control—lacks control) reflect a term’s intensity in a given domain, ranging from 0 to 1, with greater values reflecting more positive valence, stronger arousal, or greater dominance. The NRC EI, NRC WEA, and NRC VAD lexicons have been translated by their developers into over 100 languages and, although some cultural differences exist, their stability has been shown across languages (Mohammad, 2020; https://www.saifmohammad.com/).
Procedures and Analysis
Data Preprocessing
All test items were preprocessed prior to linguistic analysis. First, test item text was standardized to make items easier to process by converting all letters to lower-case, removing punctuation, and converting all numerals to words. As the tasks of this study were mostly aligned with traditional sentiment classification and analysis at the sentence (test item) level, no tokenization was employed during preprocessing. Stop words were not removed to avoid the unintentional removal of their potential contribution to the sentiment of a given test item. Likewise, no stemming technique or lemmatization was used. Given the unique structure of the PDS-ICD-11, wherein differentiating content is included in the response options of test items rather than in the test item statement itself, linguistic analysis was performed on the individual response options and then aggregated for each test item.
Linguistic Analysis
The Syuzhet (Jockers, 2015) package in R was used with multilingual NRC WEA, NRC VAD, and NRC EI lexicons to generate estimates reflecting emotional tone/intensity, sentiment, valence, arousal, and dominance for each test item within each language. Estimates were aggregated across all items within a given instrument to derive test-level estimates for each linguistic characteristic for each translation. A second set of test-level estimates were derived in a similar way after weighting the emotional tone/intensity, sentiment, valence, arousal, and dominance values by the total word count of a given translation. Descriptive data were then compared across translations at both the test- and item-level, and cross-translation estimates weighted for word count and their accompanying 95% confidence intervals (CIs) were used to compare the linguistic characteristics of the LPFS-BF 2.0 to those of the PDS-ICD-11.
Percentage agreement can be an inflated estimate of uniformity due to chance agreement. As such, consistency of each linguistic characteristic estimate across translations at the overall test level was also evaluated by calculating intraclass correlation coefficients (ICC; two-way mixed effects consistency of a single measurement [3,1]) and Krippendorff’s alpha (α) as additional measures of uniformity with corresponding 95% confidence intervals computed for each ICC and Krippendorff’s α. ICCs and Krippendorff’s α’s were computed using the psych (Revelle, 2024) and kripp.boot (Proutskova & Gruszczynski, 2007) packages in R, respectively. Code for linguistic analysis and computation of consistency metrics is available at https://osf.io/urghw.
Results
Test Level Comparisons
Overall levels of emotional tone/intensity, sentiment, valence, arousal, and dominance for the LPFS-BF 2.0 and PDS-ICD-11 and their associated means, standard deviations, and 95% CIs are reported in Table 1. Parallel estimates weighted by word count are reported in Table 2. As illustrated in Figure 1, most LPFS-BF 2.0 translations possess a relatively positive overall sentiment (upward slope), whereas all but two translations of the PDS-ICD-11 possessed a relatively negative overall sentiment (downward slope). Overall sentiment of LPFS-BF 2.0 translations ranged from −9 (Polish) to +8 (Spanish), with an average overall sentiment across translations (M = +0.77, SD = 5.10) that did not significantly differ from zero based on 95% CIs. The sentiment of each LPFS-BF 2.0 translation fell within ±1.96 SDs of the mean. The overall sentiment of PDS-ICD-11 translations (M = −18.69, SD = 17.30) was found to be significantly negative based on 95% CIs and ranged from −61 (Estonian) to +11 (Czech). After weighting estimates by word count, the average overall sentiment of the LPFS-BF 2.0 (M = 0.007, SD = 0.042) was found to be more positive than the overall sentiment of the PDS-ICD-11 (M = -0.023, SD = 0.024), which remained significantly less than zero; however, the difference in overall sentiment between the two instruments was not statistically significant based on 95% CIs.
Overall Levels of Emotional Intensity, Sentiment, Valence, Arousal, and Dominance for the LPFS-BF 2.0 and PDS-ICD-11.
Note. 1 = English, 2 = Czech, 3 = Danish, 4 = Dutch, 5 = Estonian, 6 = French, 7 = German, 8 = Italian, 9 = Norwegian, 10 = Polish, 11 = Portuguese, 12 = Spanish, 13 = Turkish.
Overall Levels of Emotional Intensity, Sentiment, Valence, Arousal, and Dominance for the LPFS-BF 2.0 and PDS-ICD-11 Weighted by Word Count.
Note. 1 = English, 2 = Czech, 3 = Danish, 4 = Dutch, 5 = Estonian, 6 = French, 7 = German, 8 = Italian, 9 = Norwegian, 10 = Polish, 11 = Portuguese, 12 = Spanish, 13 = Turkish.

Overall Sentiment of Instrument Translations.
Plotting the overall valence, arousal, and dominance of the LPFS-BF 2.0 and PDS-ICD-11 translations (Figure 2) revealed similar patterns across translations and for both instruments (i.e., relatively lower arousal and higher valence and dominance). Each LPFS-BF 2.0 translation was each within ±1.96 SDs of the respective mean for overall valence (M = 33.93, SD = 9.42), arousal (M = 26.48, SD = 7.48), and dominance (M = 31.38, SD = 9.61). Regarding the valence (M = 174.14, SD = 46.41), arousal (M = 152.80, SD = 49.32), and dominance (M = 164.84, SD = 48.58) of the PDS-ICD-11 translations, the Czech translation was 1.98, 2.05, and 2.00 standard deviations below the cross-translation mean, respectively, while all other translations were within ±1.96 SDs of the respective mean. When weighted by word count, average arousal and dominance of the LPFS-BF 2.0 and PDS-ICD-11 test items were similar, whereas the LPFS-BF 2.0 was again found to have a more positive valence (see Table 2).

Overall Valence, Arousal, and Dominance of Instrument Translations.
Similar patterns across translations of the LPFS-BF 2.0 and PDS-ICD-11 were seen for the overall emotional tone and intensity of these instruments (see Figure 3), and all estimates (aside from disgust for the LPFS-BF 2.0) were significantly different from zero based on their corresponding 95% CIs. Both instruments demonstrated relatively greater elevations in joy and trust across most languages, while relatively greater elevations in anticipation were observed across the majority of LPFS-BF 2.0 translations and relatively greater elevations in fear, sadness, and to a slightly lesser degree, anger were seen across most translations of the PDS-ICD-11. Of the LPFS-BF 2.0 translations, the emotional intensity of the Turkish LPFS-BF 2.0 was found to be notably stronger in anger (+1.98 SDs), anticipation (+2.61 SDs), disgust (+2.38 SDs), and surprise (+2.54 SDs) relative to corresponding means across translations, and the emotional tone of the Portuguese LPFS-BF 2.0 was notably stronger with regard to anger (+2.17 SDs) and sadness (+2.80 SDs). The emotional intensity of PDS-ICD-11 translations was predominantly within ±1.96 SDs of their respective means; however, the sadness of the Czech PDS-ICD-11 was found to 1.99 standard deviations below the cross-translation mean. Comparisons of cross-translation emotional tone and intensity estimates for the LPFS-BF 2.0 and PDS-ICD-11 weighted for word count revealed the emotional intensity of the PDS-ICD-11 was significantly stronger with regard to anger, disgust, fear, and sadness while the emotional intensity of the LPFS-BF 2.0 was significantly stronger with regard to joy (see Table 2).

Overall Emotional Tone and Intensity of Instrument Translations.
Percentage agreement, ICCs, and Krippendorff’s α’s reflecting overall consistency of each linguistic characteristic across translations of the LPFS-BF 2.0 and PDS-ICD-11 are reported in Table 3. Percentage agreement (exact agreement) ranged from 0% to 66.67% (Disgust) across LPFS-BF 2.0 translations and from 0% to 63.75% (Surprise) across PDS-ICD-11 translations, with no estimate crossing the common threshold for acceptability (typically >70% agreement). Poor consistency across translations was also illustrated by the computed ICCs and Krippendorff’s α’s in Table 3, none of which reached traditionally acceptable levels of agreement (ICC > 0.5, Koo & Li, 2016; α > .67, Krippendorff, 2019).
Overall Consistency of Linguistic Characteristic Across LPFS-BF 2.0 and PDS-ICD-11 Translations.
Note.aKrippendorff’s α computed using ordinal specification due to data type. Otherwise, Krippendorff’s α computed using ratio specification.
Item Level Comparisons
Levels of emotional tone/intensity, sentiment, valence, arousal, and dominance for each test item, averaged across translations, are reported in Tables 4 and 5 (see Tables S1 and S2 in the Supplemental Materials for parallel tables reporting item-level estimates weighted by word count). The emotional tone/intensity, sentiment, valence, arousal, and dominance for each test item within each language are presented as heatmap tables in the Supplemental Materials for the LPFS-BF 2.0 and PDS-ICD-11 (Supplemental Tables S3 and S4, respectively; weighted by word count in Supplemental Tables S5 and S6, respectively). The average sentiment of individual LPFS-BF 2.0 items (M = 0.064, SD = 0.600) was closely split, with items 4, 6, 9, 10, and 12 demonstrating relatively positive sentiment, items 2, 3, 5, 7, and 8 demonstrating relatively negative sentiment, and items 1 and 11 producing a balance between negative and positive sentiment. The estimated sentiment of item 12, averaged across LPFS-BF 2.0 translations, was 2.71 standard deviations above the mean (positive sentiment). Regarding average sentiment of individual PDS-ICD-11 items across translations (M = −1.34, SD = 3.06), eight of the 14 items possessed relatively negative sentiment (items 1, 4, 5, 7, 8, 11, 12, 13), while the remaining six items possessed relatively positive sentiment (items 2, 3, 6, 9, 10, 14). The estimated sentiment of item 13 was 2.53 standard deviations below the average sentiment for PDS-ICD-11 items across all translations (negative sentiment).
Emotional Intensity, Sentiment, Valence, Arousal, and Dominance for Individual Items of the LPFS-BF 2.0 Averaged Across Translations.
Emotional Intensity, Sentiment, Valence, Arousal, and Dominance for Individual Items of the PDS-ICD-11 Averaged Across Translations.
The greatest levels of arousal on the LPFS-BF 2.0 were seen for items 12 (+1.98 SDs) and 8 (+1.27 SDs), while item 3 generated the lowest arousal (−1.81 SDs). LPFS-BF 2.0 items 12 (+2.12 SDs) and 9 (+1.40 SDs) produced the highest levels of dominance, with item 3 (−1.67 SDs) producing the lowest levels. The most arousing items on the PDS-ICD-11, averaged across translations, were items 11 (+1.59 SDs), 10 (+1.49 SDs), and 4 (+1.13 SDs), while items 3 (−2.08 SDs) and 12 (−1.13 SDs) communicated the lowest levels of arousal. Similarly, items 10 (+1.60 SDs), 4 (+1.14 SDs), and 11 (+1.09 SDs) were found to be the items with the greatest dominance while items 3 (−1.76 SDs), 12 (−1.53 SDs), and 13 (−1.15 SDs) appeared to be the lowest in dominance.
Regarding emotional tone and intensity of LPFS-BF 2.0 items averaged across translations, item 2 evidenced the greatest levels of disgust (+2.41 SDs) and sadness (+2.16 SDs), item 7 was characterized by the greatest level of anger (+2.73 SDs), item 8 produced the strongest level of fear (+2.25 SDs), item 10 generated the greatest trust (+2.41 SDs), and item 12 was found to have the greatest levels of anticipation (+2.66 SDs), joy (+2.65 SDs), and surprise (+2.55 SDs). On the PDS-ICD-11, item 4 demonstrated the highest level of joy (+2.57 SDs), item 7 showed the greatest levels of trust (+2.26 SDs), item 11 was found to have the greatest levels of anticipation (+1.89 SDs), fear (+2.49 SDs), and surprise (+1.60 SDs), and item 13 was characterized by the greatest levels of anger (+2.06 SDs), disgust (+2.22 SDs), and sadness (+1.83 SDs).
Discussion
The objective of this article was to introduce and demonstrate a methodological approach that leverages NLP to quantitatively assess variability in linguistic characteristics across instrument translations and, as a secondary objective, between instruments purportedly measuring analogous latent constructs. As demonstrated by linguistically analyzing 13 translations of two measures of personality dysfunction (LPFS-BF 2.0 and PDS-ICD-11), this methodological approach appears feasible and informative, revealing multiple insights.
First, both the overall sentiment and emotional tone/intensity of instrument translations and different measures previously demonstrated to assess analogous constructs can diverge. The implications of these differences are yet to be explored but could extend to helping explain noninvariance of test translations and why ostensibly similar instruments possess dissimilar psychometrics. The introduced methodology offers an avenue for investigating these and other pertinent questions, such as which linguistic characteristics are ideal for a psychological instrument. Should test developers strive for linguistic neutrality or do tests perform better when their linguistic characteristics match their content (e.g., a hopefulness measure comprised only of items with positive overall sentiment)? Answers to these questions stand to inform test development practices and benefit both researchers and clinicians deciding between instruments.
Second, the variability in linguistic characteristics between the LPFS-BF 2.0 and PDS-ICD-11 revealed by the demonstrated methodology inspires another interesting question. These two tests make use of dissimilar structures despite assessing comparable constructs: the LPFS-BF 2.0 uses a traditional format, with each item rated using the same Likert scale, whereas the PDS-ICD-11 employs item-specific bipolar responses. Thus, at least a proportion of the linguistic differences observed between these two instruments might be due to different test structures. Past researchers have argued that a scale’s structure is a potential source of bias and can distort psychometrics (e.g., produce artificially inflated covariation within- and between-tests using similar Likert scales and anchors; Podsakoff et al., 2003). It’s reasonable to provisionally extend this argument to an instrument’s linguistic characteristics, begging the question, what role does test structure play when comparing the linguistic characteristics of different instruments and does the interaction between structure and linguistics impact a test’s psychometrics? The methodology introduced herein offers a strategy for empirically studying these questions, the answers to which might advise test developers or translators to consider different linguistic characteristics depending on an instrument’s structure.
Finally, study results showed linguistic characteristics can also profoundly vary across test items within a single instrument. Thus, the insights and questions posed above also appear applicable at the item level. Other item-specific research questions are also salient. For instance, the demonstrated methodology could be combined with item response theory methods to compare the difficulty and discrimination of semantically neutral test items, such as items 1 and 11 on the LPFS-BF 2.0, to the difficulty and discrimination of items whose sentiment matches their content, such as item 2 (overall negative sentiment; “I often think very negatively about myself.”). Significant differences would offer evidence supporting the importance of linguistic heterogeneity of test items. To this end, future studies are needed to not only illuminate the implications of item-level differences but also determine the utility—or lack thereof—in understanding these differences from both psychometric and clinical perspectives.
Beyond insights gained by examining the descriptive data derived via NLP, the current study also revealed that the consistency of linguistic characteristics across instrument translations, including translations that have been previously validated, can be exceptionally poor by conventional standards (Portney & Watkins, 2000). The relatively small number of test items examined likely lowered estimates (Lee et al., 2012), but this cannot discount the fact that the majority of estimates were remarkably low based on currently available interpretive guidelines for consistency reliability. Future efforts to establish consistency, reliability, interpretive guidelines specifically for the comparison of linguistic characteristics of instrument translations could be helpful. However, the poor uniformity across translations observed in the current study points to another need of higher priority. Multiple limitations exist when employing a single lexicon in NLP (Czarnek & Stillwell, 2022). Unfortunately, few multilingual lexicons are currently (publicly) available, limiting the utility and potential of the introduced methodology for linguistically analyzing instrument translations. As such, the development of additional robust multilingual lexicons for use in psychometric research would be an advantageous undertaking, and many approaches for doing so are available (Ghali et al., 2023).
The presented results reveal another major consideration that must be addressed before the potential of the introduced methodology can be truly realized. Negations, expressions that reverse the meaning of a statement (e.g., “not sad”), are a longstanding challenge for NLP (Morante & Blanco, 2021). The presence of negations on psychological measures has also been a longstanding psychometric challenge (Kam, 2018). Negations add complexity to a test’s linguistic characteristics and seemingly played a role in the current study (e.g., the overall positive sentiment of LPFS-BF 2.0 item 12). Different strategies for handling negation in NLP have been offered over the years (Mahany et al., 2022; Mukherjee et al., 2021), with some approaches created for specific applications (e.g., medical documents; Argüello-González et al., 2023). Although problematic in most endeavors, the nature and extent of issues created by negations for the introduced methodology requires further investigation. Of particular interest would be studies investigating whether negations in test items are deleterious, compounding, or benign to a test item’s priming effect(s). That is, does the priming effect of an item with a negation (e.g., “I am not sad”) differ from an analogous item without the negation (e.g., “I am sad”) or is the priming phenomenon evoked by individual terms regardless of context and qualifiers? Assuming the former, efforts to develop negation recognition and adjustment of algorithms specific to psychological tests would be vital.
Conclusion
The methodology introduced in this article offers a promising strategy for leveraging NLP to linguistically compare measures designed to assess analogous constructs and analyze test instrument translations to better understand translation quality. However, certain limitations remain and must be addressed before the presented methodology can be implemented with greater validity, utility, and confidence. Based on initial findings, the proposed strategies for improving the introduced methodology seem to be a worthwhile undertaking.
Supplemental Material
sj-docx-1-asm-10.1177_10731911251326371 – Supplemental material for Leveraging Artificial Intelligence to Linguistically Compare Test Translations
Supplemental material, sj-docx-1-asm-10.1177_10731911251326371 for Leveraging Artificial Intelligence to Linguistically Compare Test Translations by Adam P. Natoli in Assessment
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Preregistered
This study was not preregistered.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
