Abstract
The purpose of this study was to translate and culturally adapt the Carpal Tunnel Questionnaire to produce an equivalent Korean version. A total of 53 patients completed the Korean version of the Carpal Tunnel Questionnaire pre-operatively and 3 months after open carpal tunnel release. All 53 also completed the Korean version of the Disabilities of Arm, Shoulder, and Hand questionnaire pre-operatively and 3 months post-operatively. Reliability was measured by determining the test–retest reliability and internal consistency. Test–retest reliability was assessed using intraclass correlation coefficients and paired t-tests, and internal consistency using Cronbach’s alpha coefficients. Pearson correlation analysis was carried out on the Korean version of the Carpal Tunnel Questionnaire scores and the Korean version of the Disabilities of Arm, Shoulder, and Hand scores to assess construct validity. Responsiveness was evaluated using effect sizes and standardized response means. The reliability of the Korean version of the Carpal Tunnel Questionnaire was good. The scores in the Korean version of the Disabilities of Arm, Shoulder, and Hand strongly correlated with the scores in the Korean version of the Carpal Tunnel Questionnaire. Standardized response mean and effect size were both large for the Korean version of the Carpal Tunnel Questionnaire. The study shows that the Korean version of the Carpal Tunnel Questionnaire is a reliable, valid and responsive instrument for measuring outcomes in carpal tunnel syndrome.
Introduction
There has been a recent trend toward the use of patient-based self-report instruments rather than clinician-based instruments in clinical studies, because the former better predict functional status and the latter may not represent a patients’ view or capture the full extent of disability(Kim and Jeon, 2012). Several self-report instruments have been developed to evaluate upper extremity function (Hudak et al., 1996; Levine et al., 1993; MacDermid et al., 1998). A self-administrated questionnaire for the assessment of the severity of symptoms and functional status in carpal tunnel syndrome (CTS) was devised by Levine et al. (1993). Although this questionnaire has several names, such as the Brigham and Women’s Carpal Tunnel Questionnaire (Kim and Kim, 2012), the Boston Carpal Tunnel Syndrome Questionnaire (Hobby et al., 2005), the Levine questionnaire (Zyluk and Walaszek, 2011) or the Carpal Tunnel Syndrome Instrument (Atroshi et al., 1998), we refer to it as the Carpal Tunnel Questionnaire (CTQ) in this study.
The three important properties commonly used to evaluate patient-reported instruments are their reliability, validity and responsiveness to changes (Ozyurekoglu et al., 2006), and the original CTQ has been reported to be valid, reliable and responsive in CTS patients (Amadio et al., 1996; Levine et al., 1993).
When a self-administered questionnaire is used in different countries, cultures or languages, it must be well translated and culturally adapted to maintain the validity of its content (Beaton et al., 2000; Iwatsuki et al., 2014). Studies on reliability, validity and responsiveness of the Swedish (Atroshi et al., 1998), Japanese (Imaeda et al., 2007) and Chinese (Fok et al., 2007) versions of the CTQ have been reported, but no Korean version has been introduced. Because usage of the upper limbs is closely related to sociocultural requirements, the cross-cultural adaptation process is important for the CTQ (Kadzielski et al., 2008). The objectives of this study were to translate and culturally adapt the CTQ to produce a Korean version (the K-CTQ), and to measure its reliability, validity and responsiveness.
Methods
CTQ
The CTQ provides disease-specific measures of symptom severity, functional impairment and treatment outcomes in CTS, and consists of two scales that quantify symptom severity (CTQ-S) and functional disability (CTQ-F). The CTQ-S consists of 11 questions that address the intensity and frequency of pain, numbness, weakness and loss of dexterity. Five possible responses are offered per question, and are scored from 1 (no symptoms) to 5 (severe symptoms). Results are expressed as average scores for the 11 responses. The CTQ-F is composed of eight questions that address difficulties performing daily tasks. Again five possible responses are offered and scoring is on a 5-point scale (1 to 5, where 5 indicates greatest difficulty). Results are also presented as averages.
The adaptation process
The first stage of the adaptation process involved forward translation of the English version of the CTQ into Korean by two native Korean speakers fluent in English. One of the translators was medically trained but unaware of the purpose of the translation, whereas the other was medically trained and involved in this study. The two resulting Korean draft versions were then synthesized into a single Korean version by two translators and one expert in the field of hand surgery. This version was then given to two independent native English speakers fluent in Korean for back translation into English. One of these was a medically trained orthopaedic resident and the other was a professional translator without medical training. A pre-final Korean version was created by reviewing the above-mentioned versions with regard to linguistic and cultural quality at a meeting attended by a forward translator, a back translator, a research nurse and an expert from the field of hand surgery. Discrepancies were resolved by consensus to achieve conceptual equivalence with the original CTQ. The final version of the K-CTQ (supplementary material available online) was then developed after field-testing on 15 Korean healthy individuals and 15 Korean outpatients with CTS.
Korean version of Disabilities of Arm, Shoulder, and Hand
The Korean version of the Disabilities of Arm, Shoulder, and Hand (K-DASH) questionnaire consists of 30 multi-choice items with five possible responses per item. Twenty-one items concern degrees of difficulty when carrying out different physical activities, six concern symptoms and the remaining three address psychosocial effects (Kadzielski et al., 2008). The K-DASH is scored using a 0 to 100 scale, with higher scores indicating greater disability.
Patients
This study was approved by our Institutional Review Board and all patients provided written informed consent before enrolment. Patients who requested elective carpal tunnel release (CTR) for idiopathic, electro-diagnostically confirmed CTS were included. The exclusion criteria were an inability to complete the questionnaire because of cognitive impairment or a history of CTR in the same extremity.
Between January 2011 and December 2011, 60 consecutive patients underwent open CTR and agreed to participate in the study. One surgeon (K.J.K.) did all the surgical procedures, and local anaesthesia and a pneumatic tourniquet were used in all cases. A limited open technique was used (Bromley, 1994; Cellocco et al., 2005; Serra et al., 1997).
Seven of the 60 patients failed to attend the scheduled 3-month follow-up visit. We treated missing data using complete case analysis (Eekhout et al., 2012) and excluded the seven patients who did not complete follow-up from the study cohort, which consisted of the remaining 53 patients. The mean patient age was 59 years (range 32–86; SD 12) and eight (14%) were male.
All 53 patients completed the K-CTQ and the K-DASH pre-operatively (baseline) and at 3 months after operation (follow-up).
Statistical analysis
The paired t-test was used to compare the K-CTQ and K-DASH scores at baseline and follow-up assessments and the K-CTQ scores of the two baseline assessments conducted to determine test–retest reliability. All statistical procedures were two-sided and statistical significance was accepted for p values <0.05.
Reliability
The reliability of the K-CTQ was evaluated by analysing internal consistency and test–retest reliability.
Internal consistency was determined using the Cronbach’s alpha coefficients of inter-item correlations (Cronbach, 1946). A Cronbach’s alpha of 1.0 represents a perfect correlation among all items (Levine et al., 1993), and a Cronbach’s alpha of ≥0.7 is considered to indicate satisfactory internal consistency (Roh et al., 2011).
For test–retest reliability, we used the K-CTQ twice at the baseline assessment. The first questionnaire was completed in the outpatient clinic, and the second was completed 1 or 2 days later by telephone. Test–retest reliability was assessed using intraclass correlation coefficients (ICCs) and paired t-tests for the two administrations of the questionnaire (Kim and Kang, 2013; Rosales et al., 2002). An ICC value of >0.75 indicates that an instrument is reliable (Navarro et al., 2011).
Validity
The validity of the construct was assessed by examining correlations between K-CTQ scores and K-DASH scores. We hypothesized that the K-CTQ-S and K-CTQ-F scores would be positively correlated with the K-DASH score. Pearson correlation coefficients were used to examine construct validity because K-CTQ and K-DASH scores were continuous variables and were normally distributed. Correlation coefficients were rated as follows: strong correlation >0.6; moderate ≤0.6 and >0.3; and poor ≤0.3.
Face validity refers to the degree to which a questionnaire indeed looks as though it is an adequate reflection of the construct to be measured (Mokkink et al., 2010). We assessed face validity using the ceiling and floor effects. The floor effect occurs when the respondents provide the worst possible scores for all items, and the ceiling effect occurs when the respondents report the best possible scores (Stucki et al., 1995).
Responsiveness
Responsiveness was assessed by comparing K-CTQ scores at baseline and follow-up using standardized response means (SRMs) and effect sizes (ESs). We also measured the SRM and ES of K-DASH scores and compared these with those of the K-CTQ. SRM was calculated by dividing observed mean change by the standard deviation of the observed change, and ES was calculated by dividing observed mean change by the standard deviation of baseline scores (MacDermid and Tottenham, 2004; Stratford et al., 1996). We applied Cohen’s threshold for values of SRM and ES, namely: trivial (<0.2); small (≥0.2 <0.5); moderate (≥0.5 <0.8) or large (≥0.8) (Beaton et al., 1997). Large ES or SRM values indicate more improvement in the questionnaire scores from the baseline assessment to the follow-up assessment, which is achieved only when the intervention is successful and the questionnaire is sensitive for the disease.
Results
Cross-cultural adaptation
Although no major linguistic or cultural problems were encountered during the forward and back translations, some minor discrepancies arose due to linguistic and cultural differences. The fifth item of the CTQ-F, which reads ‘opening of jars’ represents a rotational movement of the hand. The corresponding word for jar in Korean is ‘dan-ji’, however, the shape and function of the Korean traditional ‘dan-ji’ differ from ‘jar’ as understood in Western culture. Therefore, we modified ‘opening of jars’ to ‘opening of screw-topped bottles’. In addition, the sixth item of the CTQ-F, which reads ‘household chores’, representing a variety of day-to-day tasks, has no appropriate Korean equivalent; therefore, we changed ‘household chores’ to ‘household cleaning’.
No patient found it difficult to understand the translated items, and the 53 patients responded to all items of the K-CTQ.
Assessment
Mean values and standard deviations of the K-CTQ and K-DASH at baseline and follow-up are shown in Table 1. At follow-up, the mean K-CTQ-S, K-CTQ-F and K-CTQ scores had decreased significantly from their baseline values.
Baseline and follow-up assessment of K-CTQ (K-CTQ-S and K-CTQ-F) and K-DASH.
ES: effect size; SD: standard deviation; SRM: standardized response mean.
Reliability
The internal consistency of the K-CTQ was high, as shown by the value of Cronbach’s alpha, and the alpha values of its symptom and function subscales also showed excellent reliability (Table 2). A total of 50 of the 53 patients were involved in the test–retest reliability of the K-CTQ. The mean scores of the two administrations of the K-CTQ-S, K-CTQ-F and K-CTQ were not significantly different (Table 2).
Reliabilities of the K-CTQ.
ICC: Intraclass correlation coefficient.
No statistically significant difference was observed in paired t-tests (p > 0.05).
Validity
A strong correlation was found between K-CTQ and K-DASH scores. The correlation coefficient between K-CTQ-S scores and K-DASH scores was 0.61 (p < 0.001), the correlation coefficient between K-CTQ-F scores and K-DASH scores was 0.67 (p < 0.001) and the correlation coefficient between K-CTQ scores and K-DASH scores was 0.71 (p < 0.001).
At baseline, no patient had a K-CTQ score of zero (ceiling) or a maximum score (floor). At follow-up, one patient (2%) had a K-CTQ score of zero (ceiling), but no patient had a maximum score (floor).
Responsiveness
The SRM and ES values of the various scores are shown in Table 1.
Discussion
During the development of the K-CTQ, we encountered only minor problems in translation and cultural adaptation. The K-CTQ proved to be a reliable and responsive instrument. The validity of K-CTQ was also indicated by a strong correlation with the K-DASH.
Internal consistency, assessed using Cronbach’s alpha coefficient, was high for the K-CTQ, K-CTQ-F and K-CTQ-S, as it is in the original CTQ and its Chinese (Fok et al., 2007), Swedish (Atroshi et al., 1998) and Japanese versions (Imaeda et al., 2007). Test–retest reliability was assessed using ICCs; the ICCs of the K-CTQ, the K-CTQ-S and the K-CTQ-F were high, despite the fact that retesting was carried out over the telephone. Furthermore, K-CTQ test–retest reliabilities were comparable with those of the original CTQ and its Chinese (Fok et al., 2007), Swedish (Atroshi et al., 1998) and Japanese versions (Imaeda et al., 2007). This finding implies that the K-CTQ is easily understood and answered, and that its reliability is maintained even when it is administered over the telephone.
The validity of the K-CTQ was demonstrated by the high correlation between K-CTQ and K-DASH scores. The DASH is a well-known, frequently used region-specific measure for assessing upper extremity disabilities, and has been shown to be reliable and valid in patients with upper extremity disease or trauma (Hudak et al., 1996). Furthermore, it has been translated and culturally adapted into the Korean language and provides a reliable and valid means of quantifying upper extremity problems in Koreans (Kadzielski et al., 2008).
The CTQ is a disease-specific instrument, which was originally designed to measure symptom severity and functional impairment associated with CTS (Gay et al., 2003). Therefore, after CTR, the CTQ is more responsive than the DASH, which is a region-specific instrument. The SRM and ES of the English version of the CTQ were found to be greater than those of the DASH at 3 months after CTR (Gay et al., 2003). Greensdale et al. (2004) also reported a greater value in SRM when comparing CTQ-S and DASH at 3 months after CTR. In the Spanish version, the SRM and ES values of the CTQ-S were also larger than those of the DASH at 3 months after CTR (Rosales et al., 2009). We also found that the K-CTQ had larger SRM and ES values than the K-DASH, which means that the responsiveness of the CTQ is maintained by the K-CTQ.
This study has several limitations. First, the time interval used to determine test–retest reliability was only 1 or 2 days, which is shorter than that used by others (Atroshi et al., 1998; Rosales et al., 2002). Although the optimal interval for determining test–retest reliability has not been determined, too short a time presents the risk that patients remember their answers to questions in the previous administration. Second, the second administration of test–retest reliability was carried out over the telephone. Some authors have reported that respondents tend to report better health and milder symptoms when surveys are done over the telephone rather than by mail (Brewer et al., 2004; Hays et al., 2009).
Footnotes
Conflict of interests
None declared.
Ethical approval
Authors certify that Ewha Womans University School of Medicine has approved the reporting of these cases, that all investigations were conducted in conformity with ethical principles of research and that informed consent for participating in the study was obtained.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
