Abstract
Little is known about the reliability and validity in web surveys, although this is crucial information to evaluate how accurate the results might be and/or to correct for measurement errors. In particular, there are few studies based on probability-based samples for web surveys, looking at web-specific response scales and considering the impact of having smartphone respondents. In this article, we start filling these gaps by estimating the measurement quality of sliders compared to radio button scales controlling for the device respondents used. We conducted therefore two multitrait–multimethod (MTMM) experiments in the Norwegian Citizen Panel (NCP), a probability-based online panel. Overall, we find that if smartphone respondents represent a nonnegligible part of the whole sample, offering the response options in form of a slider or a radio button scale leads to a quite similar measurement quality. This means that sliders could be used more often without harming the data quality. Besides, if there are no smartphone respondents, we find that sliders can also be used, but that the marker should be placed initially in the middle rather than on the left side. However, in practice, there is no need to shift from radio buttons to sliders since the quality is not highly improved by providing sliders.
Keywords
The increase in the number of web surveys has been accompanied by an increase in the literature studying the comparability of the data coming from web surveys with the one coming from more traditional modes of data collection, in terms of representativeness and of data quality (Cobanoglu, Warde, & Moreo, 2001; Dillman, 2000; Dillman et al., 2008; Funke, 2016; Heerwegh, 2009; Kreuter, Presser, & Tourangeau, 2008), as well as specific challenges and opportunities of web surveys (see for instance, Ilieva, Baron, & Healy, 2001).
Different studies compared typical web scales to more traditional scales. Some study drag-and-drop scales (e.g., Kunz, 2015; Sikkel, Steenbergen, & Gras, 2014), drop-down menus (Christian, Dillman, & Smyth, 2007; Couper, Tourangeau, Conrad, & Crawford, 2004; Leeuw, Hox, & Dillman, 2008; Liu & Conrad, 2016), or order-by-click scales (e.g., Revilla, Ochoa, & Loewe, 2013). In this article, we are interested in sliders scales, that is, that the respondent shall answer by moving a marker along a line to rate the item.
Previous research suggests that sliders are not the preferable form of response scales. Sikkel, Steenbergen, and Gras (2014) found that slider scales are perceived more enjoyable than traditional radio button scales by respondents but only at the first-time answering this kind of scale. The use of sliders is more cognitively demanding and hence completion time, break offs, and missing answer are increasing when sliders are used instead of radio button scales (Cook, Heath, Thompson, & Thompson, 2001; Couper et al., 2004; Funke, Reips, & Thomas, 2011; Husser & Fernandez, 2013; Roster, Lucianetti, & Albaum, 2015; Sikkel et al., 2014). Buskirk, Saunders, and Michaud (2015), in contrast, found that completion time for sliders starting at the middle or the right is shorter but causes higher item nonresponse than radio button scales. If the starting point of the slider is on the scale, then it is not clear whether a respondent indeed would choose that position and therefore does not move the slider or whether he or she refuses to answer. Sellers (2013) also found that the initial position of the slider marker has an effect on the response: Markers positioned initially at the high point of the scale yield higher scores, and those at the low point of the scale yield lower scores. Cape (2009) found that slider scales produced more accurate responses than other scales, while Couper, Tourangeau, Conrad, and Singer (2006) did not find a difference in the response distribution compared to other scales such as radio button scales. Finally, in their study, Buskirk and Andrus (2014) saw no differences between PC and smartphones regarding different aspects (recruitment, survey completion, and items), whereas Buskirk et al. (2015) results suggested that slider scales are more often preferred by mobile respondents but may not be as much as radio button scales. Overall, it is often recommended to stick to radio button scales or other more typical web scales instead of using sliders.
However, there are several reasons why it continues to be of interest to study the use of sliders. First, scales that a few years ago could seem hard to use or led to technical problems (and thereby likely to frustration and breaking off the survey) may perform much better nowadays because the capabilities of PCs and mobile devices and the experience of people in using such devices have increased vastly.
Second, as pointed out by Couper and Miller (2008, p. 831), “a key characteristic of Web surveys is their diversity.” Indeed, web surveys vary a lot at many levels: the topics, the kind of recruitment of the respondents and of samples, the optimization or not of this layout for mobile devices, and so on. It is very difficult to make general conclusions about the performance of response scales in web surveys based on the small body of existing literature.
Third, there is no previous research estimating the measurement reliability and validity, the complements of random and systematic errors, of these scales. However, having information about the measurement reliability and validity is crucial for several reasons: It helps designing new questionnaires (Revilla, Zavala-Rojas, & Saris, 2016; Saris & Gallhofer, 2014) and allows correcting for measurement error, which is a necessary step to get accurate conclusions (DeCastellarnau & Saris, 2014; Saris & Revilla, 2016).
Few multitrait–multimethod (MTMM) experiments have already been conducted in web surveys. Overall, this research suggests that the quality of web survey questions is quite high and similar to the one in more traditional modes like face-to-face (Decastellarnau & Revilla, 2017; Revilla, 2010, 2015; Revilla & Saris, 2013a, 2015). Nevertheless, the quality varies across languages, topics, and scales. Thus, this research is still too limited to make general inferences that could really inform web survey designers and/or to allow correction for measurement errors in web surveys. There is therefore a clear need of more evidence about the reliability and validity of survey questions in the case of web surveys.
The goal of this article is to start filling in these gaps by estimating the reliability, validity, and quality of sliders compared to radio button scales using two MTMM experiments implemented in the Norwegian Citizen Panel (NCP), a probability-based online panel in Norway. The NCP allows the respondents to answer the survey on their preferred device, on a PC, tablet, or smartphone. Hence, we can estimate the reliability, validity, and quality only for PC/tablet respondents (excluding smartphone respondents) and for the whole set of respondents (including smartphones) to see whether the presence of smartphone respondents affects the quality because both the smaller screens and the touch screen navigation are likely to provoke differences.
The remainder of this article is divided as follows: in the second section, we give further information about the data, and in the third section, we present the MTMM approach, the experiments, and the analyses more in details. In the forth section, we provide then the main results, and finally, in the fifth section, we conclude.
Data
Data Collection in the Norwegian Citizen Panel Wave 7
The NCP is a research-purpose Internet panel with more than 10,000 panel members. A probability-based sample of the Norwegian population from 18 to 95 years old was drawn from the Norwegian National Registry. Panel members complete an online questionnaire of about 20 min twice a year. As an incentive, a lottery on a travel gift card valued 25,000 NOK (about 2,650EUR) is included in each round. In general, the NCP does not provide explicit “Don’t Know” options, but respondents are not forced to answer whether they do not want. More information about the NCP can be found at http://digsscore.uib.no/methodology
Wave 7 took place during the month of November and the 2 first days of December 2016. After the invitation e-mail and three reminders, the overall response rate was 72%. A total of 4,689 panelists participated in this wave. Participants were divided in four subsets that received different questions. Respondents could participate using the device they wanted to answer the survey and the NCP recorded this information as a variable. Ideally, we would have analyzed separately all three types of devices: PCs, tablets, and smartphones. However, the NCP just differentiated between smartphone respondents and nonsmartphone respondents. Twenty-six percent of all panelists participated in the questionnaire using a smartphone and 74% used either a tablet or a PC (Skjervheim & Høgestøl, 2016). Comparing the sociodemographic characteristics of these two groups, we only find significant differences in gender and age: Men and young respond more on a smartphone (see Supplemental Online Material A).
The percentage of mobile respondents that did not complete the survey to such an extent that they were classified as nonrespondents was more double (8.7%) than for PC/tablet respondents (4.2%). This means that smartphone respondents were more likely to leave the survey before the end of the questionnaire. Unfortunately, we cannot test for a possible selective nonresponse bias since we do not have data of those excluded. Moreover, analyzing smartphone respondents alone would have led to very small sample sizes, and we can therefore just compare the results for PC/tablet respondents only and for all respondents together (PC/tablet and smartphones).
Method
The MTMM True Score (TS) Model
The MTMM approach (Andrews, 1984; Campbell & Fiske, 1959) consists in repeating a set of questions measuring correlated concepts of interest (called traits), using several methods, for instance, several scales (radio button, drag-and-drop, and slider). This approach has provided a large amount of information about the measurement quality, as product of reliability and validity, of survey questions for many different topics, countries, languages, and response scale formats (Alwin, 1997; Andrews, 1984; Revilla & Ochoa, 2015; Revilla, Saris, & Krosnick, 2014; Rodgers, Andrews, & Herzog, 1992; Saris & Gallhofer, 2007, 2014; Saris, Revilla, Krosnick, & Shaeffer, 2010; Scherpenzeel & Saris, 1997).
Different models have been proposed to analyze MTMM experiments but using the TS model proposed by Saris and Andrews (1991) appears to be the most attractive option because (1) this model provides better fit compared to others (Corten et al., 2002; Saris & Aalberts, 2003) and (2) it allows estimating the reliability and validity of survey questions separately, as can be seen in the system of equations below or in the graphical representation (see Supplemental Online Material B).
where Fi is the ith trait or factor, Mj is the jth method, Yij is the observed answer for the ith trait and the jth method, Tij is the TS or systematic component of the response, rij is the reliability coefficient, vij is the validity coefficient, and eij is the random error associated with Yij.
Equation (1) defines each observed variable as the sum of the associated systematic component and random errors. Equation (2) indicates that each systematic component is itself the sum of the trait component and the effect of the method used to assess it. Then, the total measurement quality is obtained by taking the product of the reliability and validity: q2 ij = r2 ij × v2 ij .
In addition, in the initial TS-MTMM model, we assume that (1) the random errors are uncorrelated with each other and with the independent variables in the different equations, (2) the traits are correlated, (3) the method factors are uncorrelated between them and with the traits, and (4) the impact of the method factor on the traits measured with a common scale is the same.
The Split-Ballot TS-MTMM
To be identified, such a TS-MTMM model usually requires including at least three correlated traits, each measured with three different methods. In order to reduce the cognitive burden of the respondents and to limit possible memory effects due to the repetitions of the same questions to the same respondents (van Meurs & Saris, 1990), Saris, Satorra, and Coenders (2004) proposed to combine the MTMM approach with a split-ballot approach: Respondents are split randomly into several groups; each split-ballot group gets a combination of two methods for a given set of three traits, instead of getting all the three methods. Two designs are possible: The two-group design which has the advantage that all respondents get Method 1 (M1) when answering the questions the first time and then half of the sample gets Method 2 (M2) and the other half Method 3 (M3) when answering the questions a second time. The MTMM model is still identified under quite general conditions when a split-ballot design is used (Saris, Satorra, & Coenders, 2004). However, in practice, nonconvergence problems and improper solutions occur frequently for the two-group design (Revilla & Saris, 2013b). In order to reduce these problems, in this study, we implemented the three-group split-ballot design: Respondents are randomly assigned to three groups: Group 1 gets Methods 1 and 2, Group 2 gets Methods 2 and 3, and Group 3 gets Methods 3 and 1.
The Two MTMM Experiments in the NCP Wave 7
Different experiments were included in Wave 7. In one of the four subsets of the sample, two three-group split-ballot MTMM experiments were implemented to evaluate the reliability and validity of a set of six traits. This set of experimental questions was answered by a subset of 1,199 panelists divided in three groups (372 in Group 1, 410 in Group 2, and 417 in Group 3) following the split-ballot design described just before. The average nonresponse rate for those questions was 1.5% and the proportion of smartphone respondents around 35%. The first experiment asked about the immigration in Norway (“qualification of immigrants” experiment), and the second one asked about the Norwegian Supreme Court (“evaluation of court” experiment). Table 1 presents the exact wording for each trait in each experiment.
The Three Requests for Answers Analyzed for Each Experiment.
Each request for answer was measured with three different methods, which are presented in Table 2. The methods are similar in both experiments. Group 1 always answered M1 first and later M2, whereas Group 2 answered M2 and later M3 and Group 3 answered M3 and later M1.
The Three Methods Analyzed: Main Characteristics and Visual Layout for PC/Tablet Respondents.
All three methods use partially labeled scales where the end points are not at all or completely. However, M1 is a rating scale with five radio buttons, whereas M2 and M3 are slider scales with underlying 51 points, that is, not visible to the respondent. In M2, the marker is initially placed in the middle of the scale, whereas in M3 it is placed on the left side.
As the sliders allow for a higher differentiation of the answers, we expect that the quality for M1 will be lower than the one for M2 and M3, taking advantage of the sliders’ properties to present a continuous scale.
As the marker is initially placed on the left side of the slider, we expect that the means and distributions will be shifted towards the left when asking the questions using M3 instead of M2 where the marker is initially place in the middle. Hence, we expect more measurement errors when the marker is initially placed on the left (more systematic errors).
Smartphone respondents got a different visual representation of the questions in order to enhance their experience. The radio button scales were vertical, and the sliders layout was changed (see Supplemental Online Material C).
Analyses and Testing
The split-ballot TS-MTMM models were estimated for each experiment using the maximum likelihood multiple-group estimation procedure of LISREL (Jöreskog & Sörbom, 1996) with the Pearson correlation matrices, means, and standard deviations as input data, for an example of the initial LISREL input, see Supplemental Online Material D. The groups correspond to the different split-ballot groups. The models are analyzed first for PC/tablet respondents only and then for all respondents together (PC/tablet and smartphones).
Before looking at the estimates, each model is carefully tested for misspecifications following the procedure proposed by Saris, Satorra, and van der Veld (2009), which is implemented in the JRule 3.0.4 software (van der Veld, Saris, & Satorra, 2009). A misspecification is defined as a deviation larger than .4 for the standardized loadings and than .1 for the causal effects and correlations (default values of the software). If a parameter is or not misspecified, it is identified by using information about the power of the test, the modification indices, and the expected parameter change.
Starting from the initial model described before, the model is corrected step by step when misspecifications are found, till an acceptable fit is achieved, for a list of the extra parameters freed and information about the fit of the final model used see Supplemental Online Material E.
Since we used a three-group split-ballot MTMM design, we expect differences in quality depending if the method is asked at the beginning or at the end of the survey. Respondents can learn how to answer using specific response formats, giving more accurate answers at the end of the survey (in this case, the quality will increase) or they can get tired of answering (in that case the quality will decrease), so even if, as a starting point, we assumed that the quality for a given method is similar in the different split-ballot groups, when necessary we allowed variations.
Results
Preliminary Analyses
Before getting into the main results, we did some preliminary analyses to test whether the means and distributions of the questions included in the MTMM analyses differ: (a) across time, that is, the same method asked at the beginning or at the end and (b) across methods, that is, two different methods at each point in time. Differences could be attributed to the split ballot design: Different respondents for the same method at different points in time and different respondents for the different methods at a given point in time, besides time and method. However, this is not the case: There are no significant differences between groups on the key demographic variables age, gender, municipal size, and education (see Supplemental Online Material F).
In order to make both types of scales, the categorical radio button scale and the continuous slider scale, comparable in terms of means and distributions, we converted both slider scales following two different approaches: (1) to compare distributions, we converted both continuous scales to categorical scales with five categories: 0–10, 11–20, 21–30, 31–40, and 41–50, and (2) to compare means, we used a typical rescaling method: Values in between a 0 and 50 range were transformed to values in between a 1 and 5 range using the following formula:
where xmin and xmax are the minimum and maximum possible rating in the specific scale for the item and xi is the item rating.
These preliminary analyses consider all respondents together (PC, tablet, and smartphone respondents). The methods for which we found significant differences (5% level) across time are presented in Table 3.
Methods for Which We Found Significant Differences Across Time.
There are often significant differences across time both in terms of distributions and means for the two slider scales, whereas for the radio button scale, only one significant difference is found (for the first request of the “evaluation of the court” experiment, in terms of means). Exploring the distributions, we find that for the sliders, less respondents selected the initial marker position when answering at the end, possibly increasing the accuracy of the answers. The significant differences across methods at a given point in time are presented in Table 4.
Significant Differences Across Methods at a Given Point in Time.
Note: “All” means that all three tests (M1 vs. M2, M1 vs. M3, and M2 vs. M3) lead to significant differences. B = beginning; E = end.
At a given point in time, significant differences in distributions and in means are generally found across methods. Thus, overall, we can conclude that the use of different scales to measure the same concepts leads to different distributions and means and that asking the same request using the same scale at different points in the survey also leads to different distributions and means for slider scales. The next question is: are reliability, validity, and quality different too?
Reliability, Validity, and Quality Estimates
Analyzing the two split-ballot MTMM experiments implemented in the NCP in the way described before, we obtained estimates of the reliability, validity, and quality of each request for an answer when asked with each of the three methods. Table 5 presents these estimates for each experiment, request for an answer and method, first for PC/tablet respondents only and then for all respondents. When the estimates vary depending on the time of the method, both are indicated: The estimates when the method is at the beginning (B) or at the end (E).
Reliability, Validity, and Quality Estimates for the Different Traits, Methods, Experiments for PC/Tablet Only or for All Devices Together.
Note. B = beginning; E = end; M1 = radio button scale; M2 = slider starting with the marker in the middle; M3 = slider with the marker starting on the left. The bold values are significant at 0.05.
First, Table 5 shows that the data quality is overall quite high in both experiments, for all requests for an answer and methods. The reliability is usually lower than the validity, sometimes quite a lot. To simplify the discussion, and because we are interested in comparing the quality across methods, we will focus for the rest of the discussion on a comparison of the averages across all three requests for an answer (column “Avg”). In order to assess the significance of the differences between the average coefficients, we conducted several two-tailed z tests for two means. Table 6 presents the significant differences between methods across devices (PC/tablets or all) and Table 7 presents the significant differences between devices across methods.
Significant Differences Between Methods of the Quality Estimates for PC/Tablet Only and for All Devices Together.
Note: “All” means that all three tests (M1 vs. M2, M1 vs. M3, and M2 vs. M3) lead to significant differences. B = beginning; E = end.
Significant Differences of Quality Estimates Between Devices Across Methods.
Since normally researchers are only asking once the questions and differences are not statistically significant, to simplify the discussion, we are mainly interested in comparing the estimates of the methods at the beginning. Considering all respondents, independently of the device they used to answer the survey, using the radio button (M1) or the slider with the marker starting at the middle point (M2) leads to a similar measurement quality, in both experiments. In addition, the slider with the marker starting on the left side (M3) leads to a similar quality estimate in the “qualification of immigrants” experiment, but not in the “evaluation of the court” experiment. Even though, differences between methods when considering all the respondents are not statistically significant.
If instead we consider only the PC/tablet respondents, there are significant divergences between methods: The slider with the marker at the middle (M2) is the one with the highest quality in both experiments. However, the order of the two others (radio button and slider with the marker on the left) differs. On the one hand, in the “qualification of immigrants” experiment, while the quality of the radio button scale decreases a little bit, the quality of the two sliders increases somehow, leading to a higher quality estimate for M2, followed by M3 and with the lowest quality for M1. On the other hand, for the “evaluation of the court” experiment, the quality increases a little for M1 and M2 whereas it stays the same for M3, increasing the differences between, on one side, M1 and M2, and on the other side, M3.
Focusing on the reliability and validity estimates may help to understand those differences. For instance, looking at the “evaluation of the court” experiment, we can see that M1 and M2 have equal reliability estimates independently of the device used, so quality differences are explained by their validity coefficients, which are much lower for M3.
Getting in the detail of how adding the smartphones respondents affects reliability and validity, we notice that in the “evaluation of the court” experiment most of the decrease is observed for the validity. Moreover, it seems that this is mainly because of M2. Furthermore, for the “qualification of immigrants” experiment, M2 also seems to be more responsive to adding smartphone respondents, but in this case the decrease in reliability is slightly bigger. Looking at Table 7, we see that the only significant difference between coefficients when adding the smartphone respondents is the one from Method 2 for the “qualification of immigrants” experiment.
Conclusions
In this article, our goal has been estimating the measurement reliability and validity of different questions asked using different scales in a probability-based online panel in Norway (the NCP). We have focused on a comparison of the quality in radio button versus slider scales and have considered the impact of having smartphone respondents and not only PC/tablet respondents since previous research on these aspects was scarce.
The data came from two split-ballot MTMM experiments implemented in the seventh wave of the NCP, which dealt with the qualification of immigrants and the evaluation of the Supreme Court. In each experiment, three different requests for an answer were asked each using three different methods: a radio button, a slider with the marker initially placed on the middle, and a slider with the marker initially placed on the left.
Overall, when considering the whole sample (i.e., all respondents, independently of the device they used to complete the survey), we have found that the measurement quality in the NCP is quite high for all traits and methods, in both experiments (.72 on average). It is quite comparable to what is found in previous research. The validity is in general higher than the reliability, sometimes quite a lot. On the contrary, we have found similar quality levels for the radio button and the slider with the marker starting on the middle, on both experiments, which differs from what we expected. When considering all the respondents the differences between the quality estimates of all three methods are not statistically significant, which implies that the three methods can be considered equally advisable.
When considering only PC and tablet respondents, more variations are observed across the methods and if the slider with the marker in the middle is the one with the highest quality in both experiments, the order between the two other methods changes across experiments. Our results show that when including smartphone respondents, the average measurement quality decreases, but differences are not statistically significant, which suggests that the measurement quality for smartphone respondents is not necessarily lower. Furthermore, considering that the differences between methods are not statistically significant when adding the smartphone respondents, we can point out that previous research not including smartphone respondents and/or not separating smartphone from PC/tablet respondents may not be generalizable to studies where the smartphone respondents represent a large part of the sample or even the whole sample.
However, we should keep in mind that these results have limits. In particular, they are only based on: (a) Norway, in which the Internet coverage is much higher than in the rest of the world; (b) one probability-based panel, the NCP, which might differ from other online panels; and (c) two specific topics, which are quite different the one form the other but might lead to different results of what other topics would do. Therefore, further research should test the robustness of the results.
In addition, because of the small sample size, we could not analyze the smartphone respondents by themselves but only deduced what the measurement reliability and validity could be based on the results of PC/tablets and of all respondents together. For a similar reason, we also could not separate PC and tablet respondents to see whether differences exist between these two device types. Further research in that direction would be interesting.
Finally, even if the research has limitations that should be kept in mind, this new study can help to set some practical recommendations on topics where little was known. If smartphone respondents represent a nonnegligible part of the whole sample, our results suggest that using sliders or radio button scales leads to an overall quite similar measurement quality for different questions and topics. In practice, it means that sliders could be used more often without harming the data quality. Besides, if there are no smartphone respondents, sliders can be used, but placing the marker initially in the middle should be preferred to placing it on the left side. However, radio buttons being often simpler to implement, we can also say that there seems to be no need to change them for sliders since the quality is not highly improved by providing sliders. Nevertheless, we should remark that we only studied here the measurement quality and that other aspects could be affected, for instance the dropout rate, item nonresponse or the overall satisfaction of the respondents with the completion of the survey.
Supplemental Material
Supplementary_Material - Measurement Reliability, Validity, and Quality of Slider Versus Radio Button Scales in an Online Probability-Based Panel in Norway
Supplementary_Material for Measurement Reliability, Validity, and Quality of Slider Versus Radio Button Scales in an Online Probability-Based Panel in Norway by Oriol J. Bosch, Melanie Revilla, Anna DeCastellarnau, and Wiebke Weber in Social Science Computer Review
Footnotes
Authors’ Note
We are thankful to the Norwegian Citizen Panel team to accept our proposal and for their help in setting up the experiment.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
