Abstract
The study aimed to identify methodological confounding factors affecting patient satisfaction survey results. The data gathered from CINAHL and PubMed databases consisted of 355 surveys published from 2006 to 2012. Linear regression and Bayesian models, with seven potential survey-related confounders together with patient age and gender as explanatory variables, were constructed. According to the linear model, up to 12% of the original variation in patient satisfaction was explained by confounding variables, not by the actual variation in satisfaction. The presence of an interviewer resulted in lower satisfaction levels, and the satisfaction results correlated negatively with the number of items in the questionnaire. According to the Bayesian model, if patients were over 60 years old and the questionnaire consisted mainly of positively phrased items, the probability of rating their experiences as very satisfied was 75%. The Bayesian and linear models endorsed each other and revealed specifically that the surveys reporting high patient satisfaction could be predicted on the basis of confounding variables. The following recommendations are given for constructing a patient satisfaction survey: use neutral rather than negatively or positively phrased items, and use enough items to increase the likelihood that the least satisfactory care components are also included in order to better enable comparisons across sporadic surveys.
Introduction
Patients’ satisfaction with health care has been used as a measure of service quality for decades, although the validity and reliability of patient satisfaction surveys are often questioned. Typically, criticism has been targeted at inadequate reporting (Sitzia, 1999), lack of comprehensive psychometric analyses (Hankins et al., 2007) and ignoring non-response bias (Boscardin and Gonzales, 2013; Sitzia and Wood, 1998). Good reliability and validity are of particular importance when the ultimate goal of research is to generalise results from case studies to a wider population. This also holds true in patient satisfaction; sporadic patient satisfaction measures lose their wider meaning if ratings cannot be compared before and after intervention, between patient populations and between settings.
Factors affecting patients’ perception of care can be classified as care provider related (e.g. organisational structures), patient related (e.g. socio-demographics) and survey related. Compared with care provider- and patient-related factors, survey-related factors have seldom been chosen as the main target of interest. This is somewhat surprising, as ‘calibration of measuring tools’, i.e. questionnaires in the case of patient satisfaction, is the only way to enable combinations of results across separate surveys (Sitzia, 1999). A psychometric comparison of instruments as such does not answer the question of why different instruments may result in different patient satisfaction levels (e.g. Camacho et al., 2009; Haggerty et al., 2011; Vrijhoef et al., 2009); empirical research, expressly from the viewpoint of confounders and meta-analysis, is required.
Survey-related factors affecting the observed level of patient satisfaction have been found to include different methods of administration (handout, postal, telephone, and Internet questionnaires, as well as face-to-face interviewing) and instrument formats (print, illustration, and voice), which may (Cohen et al., 1996; Gribble and Haupt, 2005; Shea et al., 2008) or may not (Kamo et al., 2011; Lasek et al., 1997; Peytremann-Bridevaux et al., 2006) yield contradictory results depending on the components of care included. Both the timing of the response (Jackson et al., 2001; Lin et al., 2007; Quintana et al., 2006; Stevens et al., 2006) and response rate (Mazor et al., 2002; Perneger et al., 2005) have effects on survey results and, moreover, the former affects the latter (Lin et al., 2007; Perneger et al., 2005; Saal et al., 2005). In some cases, however, the effect of response rate on the average level of patient satisfaction can be negligible (Hekkert et al., 2009; Lasek et al., 1997). Type of scale, such as Visual Analogue and Likert, wording of the rating scale, such as ‘poor’, ‘fair’, ‘good’, ‘very good’, and ‘excellent’, and the use of positively or negatively phrased items instead of neutral items may affect the observed patient satisfaction in many ways: its average level, variability and distribution, as well as construct validity of the questionnaire (Atkinson et al., 2004; Biering et al., 2006; Hendriks et al., 2001; Karthikeyan et al., 2007; Ware and Hays, 1988).
In the present study the focus was on survey-related variables, which can potentially have a confounding impact on the results of patient satisfaction surveys. To be better able to interpret and generalise the effects of survey-related variables on the observed level of patient satisfaction, two well-known patient-related factors – age and gender – were included in the analysis as ‘control’ variables, because they have been widely reported in patient satisfaction surveys and their associations with patient satisfaction have been previously highlighted. Age is the most common patient-reported factor to be associated with the level of patient satisfaction (Cohen, 1996; Hall and Dornan, 1990; Hekkert et al., 2009; Kvist et al., 2014; Rahmqvist, 2001; Tervo-Heikkinen et al., 2008), although there is no clear consensus regarding the reasons for this relationship. Age has a strong tendency to correlate with other variables affecting patient satisfaction (Quintana et al., 2006), in particular with health status (Cohen, 1996; Jaipaul and Rosenthal, 2003) and item non-response (Voutilainen et al., 2014), of which the latter relationship, a positive correlation between item non-response and age, may partly be due to unwillingness of older people to criticise health care (Bowling, 2002; Owens and Batchelor, 1996). The direction of the relationship between satisfaction and patients’ gender varies between studies, although most results suggest that men report higher levels of satisfaction than women (Elliott et al., 2012; Findik et al., 2010).
Aim
The purpose of this study was to search for potential methodological confounding factors affecting the interpretation of patient satisfaction survey results, in order to contribute to reliable comparisons across patient satisfaction surveys and their meta-analytic combinations. As this study used a meta-analytic approach, the statistical unit in the analyses was one survey dataset corresponding to one journal article. A patient satisfaction survey was defined as a survey in which explicitly patient satisfaction to care given by formal health care professionals was quantitatively measured.
Methods
Data collection
Data were systematically gathered from CINAHL (EBSCO Publishing) and PubMed (National Center for Biotechnology Information, US National Library of Medicine) databases. Sampling was restricted to surveys published between 2006 and 2012. Surveys published after 2012 were omitted because data gathering was targeted only at complete database years to ensure a repeatable sampling procedure. Unsatisfactory availability of full text concerning surveys published before 2006 restricted the data gathering to newer articles. Only CINAHL and PubMed databases were searched, due to practical reasons. Since the actual purpose was to collect a representative (systematic) sample of patient satisfaction surveys, not to gather all patient satisfaction surveys, and create a dataset with as large as possible variation in patient satisfaction scores, there were no logical reasons to expect a systematic error due to the use of only CINAHL and PubMed. The databases were last accessed on 6 March 2015. The data were gathered by one person within a limited timeframe. No need for cross-checking was seen, as only numerical unequivocal information was extracted and the quality of original studies was not assessed. Assessing the quality of studies included is a typical procedure in systematic literature reviews but not in meta-analyses, as it produces subjective elements unsuitable for statistical analyses per se.
The phrases ‘patient satisfaction’ and ‘care’ in the article’s title and/or abstract resulted in 11,685 and 6325 full-text English-language scientific peer-reviewed journal articles in CINAHL and PubMed, respectively (Figure 1). The number of duplicates was not counted when collecting the original material. As CINAHL was accessed prior to PubMed, an article found in PubMed was omitted if it had already been found in CINAHL. Abstracts of the articles were read to perform the first selection. The article was excluded if: (1) it was a review or meta-analysis, thus combining more than one statistical units; (2) it was a qualitative study, thus providing no numerical information required for the meta-analysis; (3) patient satisfaction was measured using a continuous scale, such as the visual analogue scale, not Likert-type ordinal scale; (4) the care was not given by formal health care professional(s); (5) the satisfaction did not clearly refer to a person’s own experience, but to a spouse’s, child’s or the entire family’s experience; (6) the proportion of patients <12 years old was >25%; or (7) the satisfaction referred only to one component of care, such as treatment outcome, patient education, information, medication, pain relief and caregiver’s interpersonal/technical skills. The third exclusion criterion was due to the fact that Likert-type scales are much more common in patient satisfaction surveys than continuous scales, and the comparison between the ordinal and continuous scales is not possible without losing some of the information provided by the continuous scale, as the continuous data need to be ordinalised prior to the comparison. The sixth exclusion criterion was applied because in most cases it was not possible to ensure that the child’s parents had not answered the questionnaire on behalf of the child. The seventh exclusion criterion was based on the preliminary finding that the level of patient satisfaction greatly differs across the components of care within surveys (see also Hall and Dornan, 1988). Therefore, the components would have been controlled in the meta-analysis, which in turn was impossible because, in most surveys, component-level satisfaction scores were not reported. Articles reporting patients’ evaluations concerning care quality were forwarded to the second selection, but included in the final analyses only if it became clear that the quality of care was measured with items referring to patient satisfaction. No grey literature and literature published in a language other than English were searched for.
Flowchart of data collection.
Next, the 418 articles resulting from the first selection were read thoroughly. The exclusion criteria in the second phase were: (1) no information concerning patients’ age and gender was provided (25 articles excluded); (2) the same data were reported in another article (17 articles); (3) patient satisfaction did not refer to care received, but to expectations about forthcoming care (eight articles); (4) no unadjusted (raw) values regarding satisfaction were given (six articles); (5) satisfaction was evaluated by somebody else other than the patient (three articles); (6) satisfaction was measured using a continuous scale (three articles); and (7) satisfaction referred to only one component of care (one article). The second round of selection resulted in 355 surveys.
Characteristics of data
The surveys (n = 355, Supplementary Table S1) represented all medical specialties, oncology (52 surveys) and psychiatry (38) being the most common, as well as all types of health care organisations, from large university hospitals to home care and ambulance services. Data for the original surveys were sampled from 53 different countries including 123 datasets from the United States, 25 from the United Kingdom, 21 from Canada, 20 from Sweden, 19 from China including Hong Kong and Taiwan, 16 from the Netherlands, 13 from Germany and 12 from Australia. For five surveys, the data were sampled from more than one country. The date when patients had received the care which they evaluated in the survey was given in 266 out of 355 cases (75%). The average time delay between the care experience and publication of the results was 4.2 years with a range of 3.6–5.0. Surveys were published by 208 different journals and only 14 journals provided more than three surveys.
Dependent variable
For the meta-analysis, all patient satisfaction values had to be transformed to the same scale. Typically, this is done by means of linear conversion, which is a valid solution if the original Likert scale is symmetric (balanced), i.e. the number of negative points equals the number of positive points and thus the neutral point locates in the middle of the scale. Originally symmetric Likert scales, such as 1 = highly unsatisfied, 2 = unsatisfied, 3 = neither unsatisfied nor satisfied, 4 = satisfied, 5 = highly satisfied, were linearly converted to a 0–100 scale by:
A linear conversion as per Equation (1) is not adequate when the Likert scale is asymmetric. An asymmetric scale means that the number of answer options expressing dissatisfaction is not equal to the number of options expressing satisfaction. In other words, the neutral answer option is not in the middle of the scale. For example, if the asymmetric Likert scale: 1 = poor; 2 = fair; 3 = good; 4 = very good; 5 = excellent is linearly converted to a 0–100 scale, 50 does not refer to neutral patients, but to those who evaluate their care as good, which, implicitly, is wrong and misleading (Figure 2). Therefore, originally asymmetric Likert scales were converted to a 0–100 scale by polynomial functions (Supplementary Table S2). The Likert scale structure as such, symmetric vs. asymmetric, was not treated as an explanatory variable in the meta-analysis, but the scale effect was eliminated by means of the non-linear conversion functions. Already on the basis of preliminary analysis, it was obvious that the symmetric scale would result in a higher level of satisfaction than the linearly converted asymmetric scales, irrespective of other explanatory variables tested. In general, asymmetric Likert scales are used due to the ceiling effect (Hessling et al., 2004), which refers to the situation when respondents are so satisfied that they choose mainly the most positive answer options. An asymmetric response scale diminishes the ceiling effect by ‘stretching the positive end of the scale’, which helps to distinguish smaller differences in satisfaction by increasing response variability (Ware and Hays, 1988).
On the left, relations of a symmetric and two asymmetric 5-point Likert scales to a 0–100 scale. Grey colour symbolises the amount of error generated, when an asymmetric Likert scale is linearly converted. On the right, the relationship between the approximated and actual mean ages in the 29 surveys, based on which Equation (2) to approximate the mean age from the median age was generated.
In the surveys (n = 355), patient satisfaction was measured by asking several questions concerning different components of care (n = 106), asking a single question concerning overall care (n = 72), or by using both the multiple questions and single question methods (n = 117). In 86 cases, the same Likert scale was applied for both the multiple questions and single question methods, which enabled comparison of the methods. In 68 out of the 86 cases, the single question resulted in higher patient satisfaction compared with the average value of multiple questions (83.7 ± 8.6 vs. 80.1 ± 8.9, mean ± SD; paired samples t-test, t85 = 6.93, p < 0.01). Consequently, two separate non-combinable satisfaction scores, a single question concerning overall care and arithmetic mean of many questions, were provided in 117 cases. As each survey (statistical unit) can be presented in a meta-analytic dataset only once, the choice of exclusion was between the 72 single-question surveys and 106 multiple-questions surveys. The final meta-analysis was carried out on the surveys in which patient satisfaction was measured by asking several questions (n = 106 + 117 = 283), and surveys in which patient satisfaction was measured solely by asking a single question (n = 72) were excluded.
The most common questionnaires were the Patient Satisfaction Questionnaire (Ware et al., 1976a, b) and its many modifications (14 surveys), EORTC surveys created by the European Organisation for Research and Treatment of Cancer (10 surveys), CAHPS® surveys created by the Agency for Healthcare Research and Quality (nine surveys), and the Client Satisfaction Questionnaire distributed by Tamalpais Matrix Systems, LCC (eight surveys). In 88 out of 283 cases (31%), the questionnaire used was developed for the presented study.
As the statistical unit in the statistical analysis was one survey dataset, it was not possible to compose any exclusion criteria within surveys. In practice, this meant that an average patient satisfaction on a 0–100 scale was considered as the estimate of the survey-level satisfaction and used as the dependent variable in the models. In the case of intervention studies, results from control and treatment groups were combined and one average satisfaction value weighted by the number of patients per group was calculated for each survey. Correspondingly, if satisfaction was measured more than once in a longitudinal study, results of different measurements were combined to one average value weighted by the number of patients per time point. The original plan was to extract only control groups in the case of intervention studies and baseline measurements in the case of longitudinal studies. However, while collecting the data it became clear that separate summaries of data were not necessarily reported for control and treatment groups or groups measured at different time points. This led to the question: should these ‘inadequately’ reported studies be excluded, or should data be pooled over the control and treatment groups as well as over the different time points in all studies? The latter option was chosen.
Independent variables
Altogether, nine independent variables were generated; seven of them were continuous and two were dummy variables. These variables were informed in all 283 surveys forwarded to the meta-analysis. The variables are listed below. The variables 1–4 served as control variables in the meta-analysis and they were rationalised based on statistical aspects or previous research findings concerning patient satisfaction. The variables 5–9 represented potential survey-related confounding factors.
The number of patients whose responses were included in analyses of the original publication (274, 10–1,971,632; median, range). The total number of patients was chosen as an independent variable to control the impact of study size. Weighing the dependent variable by the study size would have been another option. The proportion of eligible patients who did not respond to the survey, i.e. non-response rate summed with the proportion of patients who did respond but whose answers were excluded from analyses of the original publication (23, 0–92%). The reason why some patients were omitted was very seldom reported. The choice of non-response rate as an independent variable was rationalised by previous studies reporting that response rate has effects on patient satisfaction survey results (Mazor et al., 2002; Perneger et al., 2005). In the present meta-analysis, non-response rate was used as a control variable to which the other independent variables could be compared, to concretise their practical significance. Patients’ mean age (51 years, 12–84 years) was given in 268 out of 355 surveys. For the remaining 87 surveys, the mean age was approximated by the linear equation (Figure 2):
The proportion of female patients (55, 0–100 percent). Gender was used as a control variable in the same way as non-response rate and age. The number of points in the original Likert scale (5, 2–11). The number of questionnaire items on which the average satisfaction score was based (16, 3–153). The item positivity score [0.23, (–1)–1] calculated by:
An interview vs. self-report survey (dummy variable; 0 = self-reported, 63%). In a few cases, patients were assisted in answering the survey but the role of the assistant was not described in more detail. The assisted surveys were classified as interview surveys for the present analysis because an ‘additional’ person, an assistant or interviewer, was present both in the assisted and interview surveys when the patient answered the questionnaire. Delay between the care and survey (dummy variable; 0 = no delay in 63%, with the maximum delay nearly 4 years). In this case, no delay means that inpatients completed the survey before they left the hospital and outpatients before they left following the appointment with the health care professional.
Meta-analysis
The present meta-analysis was carried out and reported following the guidelines of the Meta-analysis of Observational Studies in Epidemiology (MOOSE) group (Stroup et al., 2000). Both linear regression and Bayesian models were constructed to ensure detection of all linear and non-linear relationships between satisfaction and explanatory variables. The linear regression model was created using IBM® SPSS® Statistics 19 for Windows. The validity of the model was evaluated by means of normality and residual evaluations, collinearity diagnostics, and the Durbin–Watson statistic to detect possible first-order autocorrelation. The number of outliers was decided on the basis of standardised residuals of the preliminary model and approximated from the normal distribution according to the traditional Chauvenet’s criterion (Kirkup, 2002). Effect sizes were reported as eta squared (Levine and Hullett, 2002). Pearson’s correlation coefficient (r) was used to demonstrate the direction of the relationship between the dependent and continuous independent variables.
In general, Bayesian methods are a family of different categories of probabilistic models, which enable probabilistic inference and allow predictions based on diverse data (Fenton and Neil, 2013). Probability is a mathematical concept that can be used to represent uncertainty. In the Bayesian modelling, uncertainty is represented as probability distributions combining conditional probabilities (Lucas et al., 2004; Nokelainen and Tirri, 2010). The conditional probability is the probability of an event under the condition of another event (Fenton and Neil, 2013; Spiegelhalter et al., 2000). In the present case, the Bayesian method was used to detect non-linear relationships between the dependent (patient satisfaction) and independent variables (confounding factors), and the likelihood of evidence (average patient satisfaction per survey) occurrence depended on events such as patients’ mean age.
The Bayesian Classification Model (BCM) was produced with B-course 2.0, developed by the Helsinki Institute of Technology (HIIT) research group, Complex Systems Computation Group CoSCo in 2013. B-course’s algorithm works simultaneously with all variables, modelling the most probable model (Myllymäki et al., 2002). BCM can be described as a list of dependencies, but it is usually more convenient to express the structure in a graphical form (Lucas et al., 2004; Nokelainen and Tirri, 2010). In the graphic network, the nodes express variables and lines indicate dependencies between them (Fenton and Neil, 2013; Myllymäki et al., 2002).
Independent variables were classified following the variable distributions (Supplementary Table S3). A three-class categorisation was chosen to facilitate the simplest model that could also indicate the existence of non-linear associations. The originally binomial variables, the presence of an interviewer and delay from the care experience to survey, were dealt with as such. Patient satisfaction was categorised into two classes based on a default prior distribution, then a survey’s likelihood to belong to either the group of ‘very satisfied’ patients (score > 78.743 on a 0–100 scale) or ‘less satisfied’ patients (≤78.743) was equal.
Results
Linear regression model
The mean patient satisfaction score pooled over the 283 surveys was 79, ranging from 23 to 97 on a 0–100 scale. Four outliers related to exceptionally low satisfaction scores were excluded from the final model, which thus included 279 satisfaction surveys. The linear model explained about 12% of the original variation in patient satisfaction (adjusted R2 = 0.12, Figure 3). Patients’ mean age, the number of questionnaire items based on which the satisfaction value was calculated, the survey type (interview vs. self-report), and the item positivity resulted in the largest effect sizes of the nine independent variables, in that order (Table 1). The relationship between satisfaction and age (r = 0.24, p < 0.01) and the relationship between satisfaction and item positivity was positive (r = 0.12, p < 0.05), signifying that the mean satisfaction level increased together with the patient’s mean age and item positivity. The relationship between satisfaction and the number questionnaire items based on which the satisfaction value was calculated was negative (r = –0.19, p < 0.01), indicating that the mean satisfaction per survey decreased when the number of items increased. The mean satisfaction was lower in the interview surveys (76.4 ± 1.0, mean ± 95% CI) compared with self-report surveys (79.0 ± 0.7; t-test, t281 = 2.17, p < 0.05). To summarise the linear model, the explanatory power of patient age almost corresponded to the explanatory power summed over all survey methodological characteristics, but each variable alone, including age, explained only a very small part of the total variation. Validity of the model was supported by the lack of multicollinearity (the variance inflation factor<1.2 in each case) or autocorrelation (Durbin–Watson 2.08). Regression residuals were normally distributed, and residual plots showed a random pattern.
Linear regression model predicting patient satisfaction on a 0–100 scale with nine explanatory variables and Bayesian Classification Model showing the probability of a survey of belonging to the group of very satisfied patients with class by class conditionings of patients’ mean age, the number of questionnaire items, and ratio between positively and negatively phrased items (item positivity) together with a combination of the selected conditionings, age conditioned to class 3 (oldest patients) and item positivity class by class. Linear regression model on patient satisfaction informed at a 0–100 scale. : unstandardised coefficient; η2: effect size as eta squared.
Bayesian Classification Model
Dependencies between patient satisfaction and explanatory variables.
Removal of the variable will weaken the model probability by the % given.
Conditioning patient satisfaction with patients’ mean age and number of questionnaire items based on which the satisfaction value was calculated, class by class (Supplementary Table S4), indicated linear associations. The association was positive with the mean age. The probability of a survey belonging to the group of very satisfied patients increased from 34% in the age group of 45.1–60 years to 65% in the age group of >60 years (Figure 3). The association was negative with the number of items based on which the satisfaction value was calculated. The probability of belonging to the group of very satisfied patients decreased from 55 to 31% when the number of items increased from ≤20 to >50 (Figure 3). A signal of non-linear association was found between the high patient satisfaction (score > 78.743) and the item positivity; the direction of the association slightly changed along the conditioning (Figure 3). The combination of the conditionings indicated that if the patients were >60 years old and the questionnaire consisted mainly of positively phrased items the probability of a survey of belonging to the group of very satisfied patients was 75% (Figure 3).
The model’s classification performance by classes is shown in Supplementary Table S5. Although the classification accuracy of the model, the estimated correctness of classification performance and its reliability, was rather low (66.78%), the model’s ability to predict the belonging of a survey to the group of very satisfied patients was good: four out of five cases (80.99%) were correct. The presence of uncertainty in the BCM was at an acceptable level. The accuracy achieved with the empirical data avoided over-fitting the model and thus allowed generalisation of the results.
Discussion
It is implicitly presumed that a patient satisfaction survey measures patients’ perceptions of care, and all other factors than the care (quality) per se are confounding variables. The explanatory power of care provider-, patient- and survey-related factors in patient satisfaction models should be as close to zero as possible. The very low original variation in patient satisfaction surveys, possible because most patients are highly satisfied in general, emphasises the need for precise control of confounding variables. From this viewpoint, fortunately, the proportion of original variation in patient satisfaction explained by the present linear model (12%) based on potential survey-related confounders can be considered rather low, albeit the explanatory power was a result of only a few independent variables. The Bayesian model endorsed the linear model and revealed specifically that the surveys reporting high patient satisfaction (score > 78.743) could be predicted on the basis of confounding variables.
The positive association between the mean satisfaction score and patients’ age indicated that older patients were more satisfied. This strongly supports previous findings from several sporadic surveys (Cohen, 1996; Hekkert et al., 2009; Kvist et al., 2014; Rahmqvist, 2001; Tervo-Heikkinen et al., 2008) and strengthens the common view that older people may be unwilling to criticise health care (Bowling, 2002; Owens and Batchelor, 1996).
The number of questionnaire items, based on which the overall satisfaction score is calculated, may associate negatively with the score at least for two reasons: (1) the higher the number of items, the greater the likelihood that the least satisfactory care components will be included; and (2) respondent fatigue, which occurs when survey participants become tired of the survey task and the quality of the data they provide begins to deteriorate (Ben-Nun, 2008). The first reason may be the more likely explanation, as respondents tired of answering questions have a tendency to use the same one or two response options, i.e. straight-line responding, or a ‘don’t know’ option when eligible, to complete the survey as quickly as possible, but they have no tendency to specifically use the most negative options (Ben-Nun, 2008). Cultural and ethnic differences in questionnaire response styles (Johnson et al., 2005; Murray-Garcia et al., 2000) may also have affected the present finding about the questionnaire length. At the survey level, response rates appear to be lower for longer than shorter versions of patient questionnaires (Rolstad et al., 2011).
The use of negatively phrased items in surveys has been proved to be associated with lower levels of scale reliability (Stewart and Frye, 2004). In this study, negatively phrased items related to lower satisfaction scores and positively phrased items to higher scores. The combination of positively and negatively phrased items may be a methodological ‘risk’ if the purpose is to measure overall patient satisfaction, as the average satisfaction score depends on which particular items are negatively/positively phrased. In other words, positive phrasing may exaggerate positive perceptions and underestimate negative perceptions, and negative phrasing vice versa. Consequently, based on the present findings, the use of neutral rather than negatively or positively phrased items could be recommended. Ensuring that the least satisfactory components are also included relates to content validity.
The detected negative effect of the presence of an interviewer on satisfaction most probably is not direct, but rather transmitted by other factors. The effect may associate with patients’ low health status, which negatively affects their satisfaction with care (Cohen, 1996; Jaipaul and Rosenthal, 2003) and, presumably, sicker patients or those who require assistance in answering the questionnaire are more likely interviewed. Although many potential biasing influences on survey responses by the mode of questionnaire administration have been suggested (Bowling, 2005), more empirical evidence of the interviewer’s role in patient satisfaction surveys is needed to offer non-speculative and generalisable explanations.
The linear and Bayesian models produced similar results due to the linearity of relationships between patient satisfaction and explanatory factors. Increasing the number of potential explanatory factors of patient satisfaction will no doubt result in a more complex model, a combination of linear and non-linear associations. Valid and reliable modelling of patient satisfaction requires studying as many explanatory factors as possible, simultaneously, which requires applying methods developed for handling big data, such as artificial neural networks (Voutilainen et al., 2014), and dealing with non-linear associations, such as using Bayesian methods (Pitkäaho et al., 2011, 2015).
Many potential confounding variables had to be omitted from the present meta-analysis, since they were not reported in the original articles comprehensively enough. In principle, it would have been possible to approach the authors of the original articles to request more detailed information but, unfortunately, it was not practically possible in the present study. Any additional knowledge regarding the effects of survey-related confounding factors on measurements of patient satisfaction is welcomed, due to the paucity of patient satisfaction studies in which survey-related factors are chosen as the main target of interest. For example, the effect of survey-related variables on patient satisfaction scores were studied only in eight out of the 355 surveys which comprised the present dataset (Supplementary Table S1).
The effect of item non-response on the patient satisfaction surveys deserves to be studied more carefully than has been done so far. As the tendency not to respond instead of giving a low score seems to be related to the patient’s age (Bowling, 2002; Owens and Batchelor, 1996; Voutilainen et al., 2014), item non-response in general could be a sign of dissatisfaction and thus a potential source of additional information. Hypothetically, (i) if people are skipping items because they are reluctant to say something negative and (ii) are willing to answer to items on which they are able to give a positive response, surveys with more missing items should show higher levels of patient satisfaction than surveys with no or fewer missing items. Item non-response should not be straightforwardly contrasted with non-response bias [see Halbesleben and Whitman (2013) for potential ways to assess non-response bias], but studied as an additional confounding factor.
In future, meta-analyses in which the respondent-level effects of confounding factors on patient satisfaction scores are used as the dependent variable need to be carried out, to avoid ecological fallacy. This in turn requires enough study material, i.e. sporadic surveys in which confounding factors have been chosen as the target of interest. The ultimate meta-analytic goal should be to include all potential confounding factors, to put the patient-, survey- and care provider-related confounders in order of importance and to reveal interactions across them.
Limitations
Two medical specialties, oncology and psychiatry, represented 25% of the data. This has to be noted as a potential limitation to the generalisability of the results. The present results concern associations across surveys, and no conclusions or generalisations within single surveys can be drawn based on them. As only a complete dataset was admissible for the meta-analysis and no statistical approximation of missing data was applied, it is most probable that the list of potential survey-related confounding factors provided in this study is not complete.
Conclusions
Based on the present results, the following recommendations are given for constructing a patient satisfaction survey: use neutral rather than negatively or positively phrased items, and use enough items to increase the likelihood that the least satisfactory care components are also included to better enable comparisons between sporadic surveys. The impact of factors other than the care per se on patient satisfaction should be controlled to improve reliability and validity of patient satisfaction surveys.
Key points for policy, practice and/or research
In this meta-analysis of patient satisfaction surveys, up to 12% of the original variation in patient satisfaction was explained by confounding variables, not by the actual variation in the satisfaction and, specifically, the surveys reporting high patient satisfaction could be predicted on the basis of confounding variables. The main confounding variables were patient age, the number of items in the questionnaire, the presence of an interviewer, and the use of negatively and/or positively phrased items instead of neutral items. The impact of factors other than the care per se on patient satisfaction should be controlled to improve reliability and validity of patient satisfaction surveys.
Footnotes
Declaration of conflicting interest
None declared.
Funding
The study was funded by VTR-funding (North Savo Hospital District, Ministry of Social Affairs and Health) for GATOH-project 2014-15.
