Abstract
Paradata are a valuable source of information to better understand the response process and to determine how the data collection processes can be improved. Currently, a debate is ongoing in academia and in the professional associations of market, opinion, and social research concerning the need to inform respondents about the collection and use of their web paradata. Unfortunately, research is lacking on the design of informed consent for web paradata use and its possible consequences for respondents’ willingness to share their paradata, completion of a survey, and response behavior. The aim of this study is to provide guidance to the survey practitioner on the optimal design of informed consent for web paradata use based on experimental evidence.
Introduction
In survey research, the term paradata refers to information that describes the process of survey data collection. Paradata are a valuable source of information to better understand what happens during the response process and how the data collection process can be improved (Couper, 2000; Kreuter, 2013). In web surveys, paradata are usually collected automatically as a by-product of computer-assisted data collection. They include, for example, the device and browser that respondents used to fill out a questionnaire, the way they navigate through questions by going back and forth, changing given answers, interrupting the survey for a certain time, and so on (Callegaro, 2013; McClain et al., 2019).
Currently, a debate is ongoing on the need to inform respondents about the collection and use of their paradata, both in academia (e.g., Couper, 2017; Felderer & Blom, 2019) and in the professional associations of market, opinion, and social research (ADM e.V., ASI e.V., BVM e.V., & DGOF e.V., 2007; ESOMAR/GRBN, 2015, 2017). This debate is driven by the development of paradata scripts that make ever new paradata accessible (Heerwegh, 2003; Kaczmirek & Neubarth, 2007; Schlosser & Höhne, 2018), and by changes in data protection regulations (e.g., General Data Protection Regulation [GDPR] in the EU). To date, however, few empirical studies have been carried out on the design of informed consent for web paradata use and its possible consequences on respondents’ willingness to share their paradata, completion of a survey, and response behavior.
In general, two types of informed consent can be distinguished. Implicit consent means that respondents are informed of the intention to collect and use their web paradata; they are not asked for consent, because implicit consent is assumed when they participate in the web survey. Explicit consent means that respondents must actively agree (through an opt-in procedure) to the collection and use of their web paradata, for example, by clicking a button that indicates consent, or they must actively disagree (through an opt-out procedure), for example, by checking a box that indicates their withdrawal of consent. To the best of our knowledge, no studies have been conducted to examine the effects of using different consent procedures for web paradata. However, we know from related research on linking survey data to administrative records that opt-out consent procedures can lead to higher consent rates compared with opt-in procedures (Sakshaug et al., 2016). With regard to other design characteristics, previous studies have shown that a clear description of the type of web paradata collected and the purposes for which they are used generally helps to promote a positive attitude toward the collection and use of web paradata (Couper & Singer, 2013; Kunz & Gummer, 2019). Moreover, asking for web paradata consent at the beginning of the questionnaire rather than at the end seems to build trust, and therefore, results in higher consent rates (Couper & Singer, 2013; Sattelberger, 2015). This finding is consistent with previous findings on informed consent for linking survey data to administrative records (Sakshaug et al., 2013, 2019; Sakshaug & Vicari, 2017). However, transparency regarding the collection and use of web paradata in terms of asking for web paradata consent at the beginning of the questionnaire also may heighten respondents’ sense of surveillance. It is already known that “explicit cues reminding people that a researcher is ‘watching’” may lead respondents to change their response behavior (Connors et al., 2019, p. 187). Consequently, respondents’ knowledge of the collection and use of their web paradata could lead to greater response engagement when answering subsequent survey questions. Related to this, a previous study on informed consent for web paradata found a greater number of mouse clicks when the consent question was asked at the beginning rather than at the end of a questionnaire, with other indicators of response quality (e.g., item nonresponse, straightlining, extremity) remaining unaffected (Sattelberger, 2015). In addition, respondents may start the survey but decide not to complete it, resulting in breakoff. One explanation for the reduced willingness to start and continue with a survey could be, among others, concerns about the collection and use of web paradata. However, this effect has not been investigated sufficiently so far, and the previous results are mixed (Couper & Singer, 2013; Sattelberger, 2015).
In summary, little experimental research exists on the impact of different ways of informing and soliciting respondents’ consent for web paradata use, nor on how information and consent for web paradata use affect survey completion and other aspects of response behavior. Our study aims to fill this research gap by addressing the following research questions:
Is consent for web paradata use affected by the timing and procedure by which respondents are asked for their consent?
Are there negative effects on respondents’ willingness to complete a survey?
Does the timing and consent procedure influence response behavior?
Data and methods
Sample
We conducted a consent experiment in October 2018 as part of a web survey with a sample drawn from a German non-probability online access panel. A quota sample was based on gender, age, and education. We invited a total of 5,563 active panel members, 824 of whom were screened out due to age restrictions of 18–69 years or because the quotas were full. The participation rate was 92% (n = 4,371; American Association for Public Opinion Research [AAPOR], 2016), and the overall breakoff rate was 5% (n = 238; Callegaro & DiSogra, 2008). Fifty percent of the respondents were female, the average age was 45 years, and 33% had a (subject-related) higher education entrance qualification.
Questionnaire
The survey on “Politics and Voting Behavior” included questions on, among others, party identification, political knowledge, and policy preferences. The questionnaire consisted of three initial quota questions, followed by 31 web pages of core questionnaire questions, and a few questions at the end that asked respondents to assess the survey. On average, it took respondents 22 min (Mdn = 19) to complete the entire questionnaire. Apart from the mandatory quota questions, respondents were able to skip all the remaining questions without giving a substantive answer. We used a responsive questionnaire design where the layout of the questionnaire dynamically adapts to different screen sizes. 25% of respondents completed the survey on a smartphone and 7% on a tablet. We included the Embedded Client Side Paradata (ECSP; Schlosser & Höhne, 2018) script on each web page of the survey to collect client-side paradata on, among others, time stamps, mouse clicks, and survey focus events.
Experimental design
All respondents received the same information on the collection and use of web paradata, including some example web paradata to help them better understand what paradata are and for what purposes they can be collected. The consent experiment was a full factorial 2 × 3 between-subjects design of consent timing (i.e., “beginning” = first page of the core questionnaire vs “end” = very last page of the questionnaire) and consent procedure (i.e., implicit consent vs explicit opt-in consent vs explicit opt-out consent; see Figure 1).

Wording of the web paradata information and consent procedures (translated from German).
Measures
Paradata consent and survey completion
We calculated the consent rates as the number of respondents who consented explicitly to the use of their web paradata among all sample cases (i.e., who successfully answered the quota questions, including completes, partials, and breakoffs).
We based the completion rates on the number of respondents who completed the survey (i.e., who answered at least 50% of all the survey questions) among all sample cases (i.e., who successfully answered the quota questions, including completes, partials, and breakoffs).
Response engagement
To assess response engagement (Baker et al., 2010), we included several response quality indicators (i.e., item nonresponse, degree of differentiation, length of open-ended responses) and paradata indicators (i.e., response times, number of mouse clicks, browser window switching). We excluded respondents who explicitly opposed the use of their web paradata from the analysis of these indicators. Moreover, we restricted calculations to the first 10 pages of the core questionnaire, because we assumed that potential influences on response behavior are strongest shortly after the web paradata information has been provided. 1
We calculated item nonresponse rates by counting the number of questions with missing values; in the case of open-ended questions, this measurement also included don’t know responses (e.g., nothing I can say about it) and nonsense responses (e.g., xxx, dsgjgkh).
To calculate the remaining indicators, we excluded cases with one or more missing values to questions asked on the first 10 pages of the core questionnaire (n = 1,173). 2
We calculated the degree of differentiation using five rating scales consisting of either 3, 4, 6, or 10 items. We used rho defined as
We measured the length of open-ended responses to three open-ended questions by automatically counting the number of characters using the Stata string function strlen.
We measured response times at the page-level (in milliseconds) and added them up for the first 10 pages of the core questionnaire. We treated cases with response times smaller or larger than the average response time ± 2SD as outliers and excluded them from the analysis (indicated in seconds).
We counted the number of mouse clicks at the page-level and added them up for the first 10 pages of the core questionnaire.
We measured window switching at the page-level by counting the frequency with which respondents left the browser window that hosted the web survey while they were processing the first 10 pages of the core questionnaire. We treated cases with window switching frequencies smaller or larger than the average number of window switching ± 2SD as outliers and excluded them from the analysis.
Results
Paradata consent
The overall consent rates were very high, ranging from 93.2% to 96.8% depending on the experimental conditions (see Table 1). We performed a Pearson’s chi-square test of independence based on experimental conditions that included an explicit opt-in/opt-out consent procedure to examine the effects of consent timing and consent procedure on consent rates. The overall chi-square test was significant, but when we examined consent rates by device type, we found significant differences for respondents using a smartphone, but not for respondents using a desktop device (i.e., a desktop PC, laptop, or tablet; also see Figure 2). With respect to the smartphone respondents, we found the highest consent rate when we asked for consent at the beginning of the questionnaire by using an opt-in procedure. Furthermore, an opt-in procedure led to higher consent rates (beginning: 98.0%, end: 93.8%) than an opt-out procedure (beginning: 89.3%, end: 85.4%), irrespective of the consent timing. Thus, the consent procedure seemed to be a more decisive factor than the consent timing.
Consent rates (left) and completion rates (right) by experimental condition (in %), overall and separated by device.
Note. Pearson’s χ²-test of independence with df in parentheses and pair-wise comparisons (Bonferroni correction), with superscripts (a–f) indicating a significant difference (p < .05 or less) between any two of the six (four) conditions.

Consent rates by experimental condition (in %), separated by device (error bars represent the 95% confidence intervals).
Survey completion
Overall completion rates were high and ranged between 93.4% and 96.0% (see Table 1). A Pearson’s chi-square test of independence did not reveal a significant difference, depending on experimental conditions. Separate analyses by device showed a small but significant effect for desktop respondents, but not for smartphone respondents. Desktop respondents who were asked for implicit consent at the end of the questionnaire had a significantly lower completion rate (92.9%) compared with those who were asked for implicit consent at the beginning (96.7%). Overall, percentages suggested that both an implicit and explicit consent procedure at the beginning of a questionnaire did not entail an increased risk of premature survey termination.
Response engagement
To examine the various indicators of response engagement, we computed ordinary least squares (OLS) regressions for each of our continuous dependent variables (i.e., degree of differentiation, length of open-ended responses, response times, mouse clicks), and marginalized zero-inflated Poisson (MZIP) regressions (Cummings & Hardin, 2019) for each of our count data variables that had an excess of zero counts (i.e., item nonresponse, window switching). The regression models included controls for key respondent characteristics (i.e., gender, age, education, and device type), which we related to computer/Internet affinity and response behavior.
In our models, we selected the experimental condition with an implicit consent at the end of the questionnaire as a reference category because, in this case, respondents completed the questionnaire without their knowledge of the collection and use of their web paradata. We found significant effects of the experimental conditions for only two dependent variables: the degree of differentiation and the number of mouse clicks (see Table 2). An explicit opt-out consent at the beginning of the questionnaire encouraged a slightly higher degree of differentiation, F(9, N = 2,255) = 10.59, p = .000, whereas an implicit consent at the beginning of the questionnaire led to a significantly higher number of mouse clicks, F(9, N = 2,255) = 13.31, p = .000, while respondents were answering the questions on the first 10 pages of the core questionnaire. Separate analyses by device type revealed that a significant effect regarding the degree of differentiation was found only for smartphone respondents, whereas a significant effect regarding the number of mouse clicks was found only for desktop respondents (see Appendix 1 Table 3). 3
Results from OLS and MZIP regression models predicting response engagement.
Note. For better readability, we did not report the coefficients for the control variables (results available upon request). OLS = ordinary least squares; MZIP = marginalized zero-inflated Poisson, b = unstandardized coefficients from OLS regression; SE = standard error; AIC = Akaike information criterion; BIC = Bayesian information criterion; na = not applicable.
b = coefficients from MZIP regression.
p < .05. **p < .01. ***p < .001.
Summary and conclusion
In this study, we examined different designs of informed consent for web paradata use and its possible consequences on respondents’ willingness to share their paradata, completion of a survey, and response behavior. Web paradata consent was generally high for our online access panelists. However, especially among smartphone respondents, when and how consent was obtained seemed crucial. With regard to consent rates, our results confirmed the advantage of asking for web paradata consent at the beginning of the survey rather than at the end. Furthermore, regardless of consent timing, an opt-in procedure was more advantageous than an opt-out procedure. Survey completion rates were high in all experimental conditions, which suggested that both implicit and explicit consent for web paradata use at the beginning of a survey did not have a negative impact on respondents’ willingness to complete a survey. In general, we did not find effects regarding various indicators of response engagement. We, therefore, concluded that response behavior is unaffected by whether and how consent for web paradata use is obtained―even for survey questions that are answered immediately after.
This study is not without limitations that should be addressed by future research. We conducted our consent experiment with online access panelists. Online access panels have become especially prevalent in market and social research, as more and more surveys have moved online (Baker et al., 2013; Comley & Beaumont, 2011). However, it should be noted that the members of non-probability panels self-select into panels and, therefore, are likely to be particularly survey-affine individuals with a high motivation to participate, and driven by intrinsic interests, financial incentives, or other factors (Baker et al., 2010; Boyle et al., 2017; Göritz, 2014). This circumstance could be an explanation for the high web paradata consent rates and survey completion rates that we observed across all experimental conditions in our experiment. To explore the generalizability of our findings, we encourage replication with probability-based panels and surveys with less motivated participants.
Implications for practice
We conducted our study from a survey methodologist’s perspective with respect to how consent for paradata use in web surveys should be designed to positively influence respondents’ willingness to share their paradata without producing a negative impact on their survey completion and response behavior. As argued previously, this approach was motivated by a lack of research and guidance on this topic, and an ongoing debate that involves academic and commercial actors. We did not aim to answer the general question of whether informed consent for paradata use in web surveys should be obtained or under what circumstances (e.g., for all paradata, for certain types of paradata alone or in combination with survey data, for certain analytical purposes, etc.). This general question is a more complex issue that involves different perspectives. An answer depends primarily on legal requirements and ethical considerations, but should not be made without considering the consequences for survey research. In our view, answering this question is the task of a group of scientists and practitioners with different disciplinary backgrounds, following the example of the AAPOR task force reports.
The general guidelines on best practices and standards issued by professional associations are necessarily rather broad and cover various topics. This approach seems to us to be the consequence of the need to serve the various forms that a data collection project can take in practice. To avoid inflating general guidelines with dense discussions on specific or niche topics, we argue for the publication of supplementary task force reports on specific topics. Our study shows what a survey methodologist’s contribution to such a report could be. In our view, it would be beneficial for several reasons to have supplementary reports on specific topics as additions to the more general guidelines.
First, professional associations could more easily take up emerging topics without having to completely revise the existing guidelines, which would allow for a faster production of professional guidelines. Ultimately, this approach could provide regularly updated recommendations for the profession to better keep pace with scientific innovation and practical needs.
Second, a clear definition of topics for supplementary reports would enable the recruitment of experts on the subject to form task forces, which would improve the debate and increase the quality of guidance provided for the profession.
Third, the formation of task forces with different disciplinary backgrounds (e.g., law, survey methodology) would help to develop solutions that, on the one hand, could deal with a topic in a holistic way and, on the other hand, would be applicable in practice. In our case of informed consent for paradata use in web surveys, the question of whether informed consent should be obtained always includes the concern with how this can be done ideally.
Footnotes
Appendix 1
Results from OLS and MZIP regression models predicting response engagement, overall and separated by device—based on all 27 pages of the core questionnaire.
| Dependent variable: | Item nonresponse a | Degree of differentiation | Response length | Response times | Mouse clicks | Window switching a |
|---|---|---|---|---|---|---|
| b (SE) | b (SE) | b (SE) | b (SE) | b (SE) | b (SE) | |
|
|
||||||
| Informed consent: Implicit/end | ref. | ref. | ref. | ref. | ref. | ref. |
| Implicit/beginning | –0.095 | –0.004 | 0.790 | 19.676 | 2.153* | –0.190* |
| (0.086) | (0.005) | (4.729) | (21.449) | (1.090) | (0.092) | |
| Explicit opt-in/beginning | 0.043 | 0.000 | 2.616 | 11.041 | 1.564 | 0.048 |
| (0.087) | (0.005) | (4.823) | (21.876) | (1.111) | (0.091) | |
| Explicit opt-out/beginning | –0.187* | 0.007 | –2.340 | –2.504 | 1.487 | 0.099 |
| (0.093) | (0.005) | (4.913) | (22.285) | (1.132) | (0.089) | |
| (Constant) | 0.180 | 0.535*** | 68.518*** | 739.542*** | 148.811*** | 0.832*** |
| (0.095) | (0.007) | (6.386) | (28.966) | (1.472) | (0.081) | |
| Adjusted R 2 | na | .04 | .02 | .03 | .08 | na |
| F statistic | na | 9.94(9)*** | 5.66(9)*** | 8.60(9)*** | 19.72(9)*** | na |
| AIC | 7,211.07 | –4,197.00 | 22,210.81 | 28,059.32 | 16,533.29 | 7,511.75 |
| BIC | 7,294.60 | –4,141.32 | 22,266.49 | 28,115.00 | 16,588.96 | 7,589.69 |
| N | 2,882 | 1,934 | 1,934 | 1,934 | 1,934 | 1,934 |
|
|
||||||
| Informed consent: Implicit/end | ref. | ref. | ref. | ref. | ref. | ref. |
| Implicit/beginning | –0.039 | –0.005 | –1.826 | 16.479 | 2.458 | –0.211* |
| (0.095) | (0.006) | (5.546) | (25.462) | (1.276) | (0.104) | |
| Explicit opt-in/beginning | 0.022 | –0.003 | 1.503 | –2.947 | 1.078 | 0.005 |
| (0.096) | (0.006) | (5.745) | (26.375) | (1.321) | (0.104) | |
| Explicit opt-out/beginning | –0.116 | 0.002 | –3.289 | –12.680 | 1.637 | 0.085 |
| (0.103) | (0.006) | (5.864) | (26.921) | (1.349) | (0.102) | |
| (Constant) | –0.206 | 0.547*** | 69.504*** | 777.399*** | 151.403*** | 0.825*** |
| (0.113) | (0.009) | (7.907) | (36.305) | (1.819) | (0.090) | |
| Adjusted R 2 | na | .04 | .02 | .03 | .09 | na |
| F statistic | na | 7.31(8)*** | 5.17(8)*** | 6.01(8)*** | 17.53(8)*** | na |
| AIC | 5,342.28 | –2,999.95 | 16,076.94 | 20,342.06 | 11,971.14 | 5,123.51 |
| BIC | 5,416.06 | –2,952.77 | 16,124.12 | 20,389.24 | 12,018.32 | 5,191.66 |
| N | 2,156 | 1,397 | 1,397 | 1,397 | 1,397 | 1,397 |
|
|
||||||
| Informed consent: Implicit/end | ref. | ref. | ref. | ref. | ref. | ref. |
| Implicit/beginning | –0.323 | –0.001 | 7.946 | 22.338 | 0.776 | –0.099 |
| (0.194) | (0.010) | (9.063) | (39.753) | (2.096) | (0.198) | |
| Explicit opt-in/beginning | 0.029 | 0.011 | 5.855 | 39.758 | 2.228 | 0.206 |
| (0.202) | (0.010) | (8.889) | (38.990) | (2.055) | (0.189) | |
| Explicit opt-out/beginning | –0.417* | 0.020* | 0.819 | 17.329 | 0.669 | 0.146 |
| (0.210) | (0.010) | (9.013) | (39.534) | (2.084) | (0.189) | |
| (Constant) | 1.020*** | 0.522*** | 66.146*** | 713.065*** | 146.374*** | 0.325 |
| (0.178) | (0.011) | (10.247) | (44.942) | (2.369) | (0.181) | |
| Adjusted R 2 | na | .06 | .02 | .06 | .05 | na |
| F statistic | na | 4.91(8)*** | 2.62(8)* | 4.92(8)*** | 4.55(8)*** | na |
| AIC | 1,749.59 | –1,190.44 | 6,136.04 | 7,723.89 | 4,563.33 | 1,711.18 |
| BIC | 1,809.23 | –1,151.87 | 6,174.61 | 7,762.46 | 4,601.91 | 1,766.90 |
| N | 726 | 537 | 537 | 537 | 537 | 537 |
Note. For better readability, we did not report the coefficients for the control variables (results available upon request). OLS = ordinary least squares; MZIP = marginalized zero-inflated Poisson; b = unstandardized coefficients from OLS regression; SE = standard error; AIC = Akaike information criterion; BIC = Bayesian information criterion; na = not applicable.
b = coefficients from MZIP regression.
p < .05. **p < .01. ***p < .001.
Data availability
The dataset generated and analyzed during the current study is available on request from the corresponding author.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
