Abstract
We investigate the reliability of data from the Wage Indicator (WI), the largest online survey on earnings and working conditions. Comparing WI to nationally representative data sources for 17 countries reveals that participants of WI are not likely to have been representatively drawn from the respective populations. Previous literature has proposed to utilize weights based on inverse propensity scores, but this procedure was shown to leave reweighted WI samples different from the benchmark nationally representative data. We propose a novel procedure, building on covariate balancing propensity score, which achieves complete reweighting of the WI data, making it able to replicate the structure of nationally representative samples on observable characteristics. While rebalancing assures the match between WI and representative benchmark data sources, we show that the wage schedules remain different for a large group of countries. Using the example of a Mincerian wage regression, we find that in more than a third of the cases, our proposed novel reweighting assures that estimates obtained on WI data are not biased relative to nationally representative data. However, in the remaining 60 percent of the analyzed 95 data sets, systematic differences in the estimated coefficients of the Mincerian wage regression between WI and nationally representative data persist even after reweighting. We provide some intuition about the reasons behind these biases. Notably, objective factors such as access to the Internet or richness appear to matter, but self-selection (on unobservable characteristics) among WI participants appears to constitute an important source of bias.
Introduction
There is no perfect data or data source. The lack of coverage or limited access to data puts boundaries not only on the development of knowledge but notably also on policy advice. The representative data are often dificult to access, sometimes they do not even exist (especially in the case of developing economies, societies in political transition, or with constraints on democracy). Even when data have actually been collected, they may miss the focus of policy and research relevant areas, hence the design of questions as well as the array of covered topics may fall short of the needs of the scientific community.
In response to these shortcomings, community of scholars have developed numerous projects of online surveys, 1 which are often distinguished by free access, thoughtfulness in design, and comprehensive coverage of topics. However, previous research suggests that data collected in online surveys are likely to suffer from the lack of representativeness, which may lead to a bias in estimated relationships (e.g., Evans and Mathur 2005; Granello and Wheaton 2004; Steinmetz and Tijdens 2009; Steinmetz et al. 2014; Steinmetz, Tijdens, and de Pedraza 2009; Valliant and Dever 2011). Defending the quality of the data from online surveys, Heiervang and Goodman (2011) argue that if a participating population is large enough, the problem of representativeness may be overstated. Online surveys make it possible to substantially increase the quality of the collected data through adequate design of questionnaires and complete control over its administration to respondents (Braunsberger, Wybenga, and Gates 2007). Moreover, low response rate need not be an issue, so long as it is fairly random. 2 Strengthened by these insights, researchers often rely on data from the online surveys. 3
One of the largest and the most popular online survey programs is the Wage Indicator (WI). The advantages of using such a database are manifold. It covers 96 countries, including some for which alternative sources are unavailable or nearly impossible to obtain. 4 Importantly, the WI provides information not only on wages, human capital, and demographic characteristics but also on a wide range of topics related to job and life satisfaction, work–life balance, and health, which makes it a unique data source for a variety of economic, sociological, political science, and possibly even psychological studies. Finally, in some countries, sample sizes in WI are indeed large comparable to those in nationally representative databases. These features made the WI an attractive alternative for researchers. For 2015 alone, WI website lists 61 scientific publications and policy reports that rely on the WI data. 5
Given the popularity of WI data, its reliability is of paramount importance. After comparing WI data from one selected wave with data from representative surveys for Germany and the Netherlands, Steinmetz et al. (2009) show that coefficients estimated on WI data are biased relative to the representative samples. Our objective is threefold. First, we provide a novel approach to constructing balancing weights. Unlike the procedure suggested by Steinmetz et al., we propose to use weights derived from propensity score matching rather than propensity scores per se. We demonstrate substantial gains in the balancing of the WI data relative to representative samples even with a fairly narrow set of matching covariates. Such gains owe to balancing based on relative rather than absolute frequencies. Second, using this novel approach, we are able to verify the claim that estimates from WI data are biased relative to representative samples. Unlike earlier literature, we provide these weights for a large collection of countries (in total 95 samples from 17 countries 6 —industrialized and developing alike). As benchmark for comparisons, we utilize a large collection of the nationally representative samples collected from a variety of sources: labor force surveys (LFS), household budget surveys (HBS), structure of earnings surveys, the International Social Survey Program (ISSP), and other dedicated national surveys. Our procedure for balancing the WI sample properties has two major advantages: (i) all the relevant information is utilized for estimating weights on each WI observation and (ii) once the weights are estimated, only WI data need to be adjusted to balance samples. Third, multiple years and data sources for some of the countries allow to identify the patterns of similarity between WI and representative samples.
Studies on life conditions often treat salary as a reference point and a key variable in the analysis. Indeed, level of salary may be an important indicator of life situation and it is often highly correlated with the assessment of the quality of the job and life in general. By and large, the results of our exercise demonstrate that WI data differ substantially from representative samples, that is, WI data on wages do not appear to represent the underlying population. Among 95 analyzed data sets, only 11 yield mean hourly wages similar to the benchmark representative samples. We were able to successfully balance the sample structure in virtually all of the analyzed WI samples and to reproduce demography and human capital endowment of the representative samples. Despite the balancing, majority of WI samples still yield wage distributions different from the representative samples, which puts in question the general reliability of the indicators from WI.
This article is structured as follows. First, we present insights from earlier studies on the reliability of the online surveys in the second section. The third section describes in detail the methodology with a particular reference to the construction of weights, while, in the fourth section, we present the main characteristics of the data sources employed in this study. We report the results in the fifth section, whereas in the concluding section, we discuss the implications and limitations of our study.
Insights From the Literature
Growing popularity of Internet data in labor market and life quality research spurred a wave of studies on the quality of these sources. This literature broadly falls into two categories. First, some studies discuss the advantages of collecting data online relative to alternative sources, which usually include references to privacy, costs, and timeliness of Internet data. However, an equally large literature discusses the methodological challenges that surround the use of data collected using online surveys.
The advantages of using Internet-collected data over alternative data sets have been discussed by several authors. First, collecting data via Internet is much cheaper and less time-consuming than obtaining similar data via field questionnaires on a representative sample of the population (Shannon and Bradshaw 2002; Wright 2005). Often, respondents in online surveys are volunteers, which makes these surveys almost costless (Horton, Rand, and Zeckhauser 2011). The reduction in costs creates the opportunity to collect data on a much larger range of topics, which are better adjusted to the needs of the researcher (Boelhouwer and Bijl 2015). Moreover, to some extent, it is possible to include populations that would be nearly impossible to reach via traditional surveys, thus the coverage can be much better (Lefever, Dal, and Matthiasdottir 2007). Finally, online tools reduce substantially the cost of collecting data, hence making more research possible within the same budget constraints.
In addition to the cost dimension, Internet-based surveys have another advantage: They provide a higher sense of anonymity for the participant than the presence of government oficials or professional interviewers (Braunsberger et al. 2007; Granello and Wheaton 2004). This feature is certainly desirable when dealing with topics that might otherwise be taboo or with information that the respondents might be reluctant to provide in the presence of government representatives or other household members (e.g., unregistered employment or earnings from illegal activities). On the flip side, the desire to preserve anonymity might result in misreporting individual characteristics, such as age or gender, even when substantial questions are later answered truthfully (Akbulut 2015).
Finally, using Internet for data collection has the potential to overcome some of the shortcomings of oficial data. Microlevel databases are collected on regularly defined periods, usually quarterly or annually. As a consequence, policy effects might only be visible after some time. Internet databases, on the contrary, are updated in real time. More importantly, oficial regulations usually place constraints on the information that can be disclosed in traditional sources. In the case of the European Union LFS, for example, the application of anonymization rules eliminated wage data from the distributed samples (income deciles are reported instead). The same anonymization process was applied to some individual characteristics such as age.
On the other side of the discussion, opponents to the use of Internet-based survey data question the quality of such data. One of the most frequently discussed issues is participation in online surveys. Since in most web-based surveys (those provided by websites or online forums) it is impossible to specify the size of the population that was able to take part in the survey, analyses of response rate or structure of the sample in comparison to total sample that was aware of the survey are also often impossible (Schleyer and Forrest 2000). Granello and Wheaton (2004) propose to deal with this problem applying a probability sampling design online. The procedure implies identifying the target population and sending surveys by e-mail only to a randomly chosen sample of potential respondents. Their proposal is only applicable to cases where the target population has been identified and e-mail addresses were collected. An alternative approach consists of using social media to spread the survey and control for the number of respondents who opened it (Ramo and Prochaska 2012). Finally, Fleming and Bowden (2009) propose to refer to the total number of visits on the website through which the survey is distributed. These measures are far from perfect, but they allow a comparison of response rates in traditional surveys to participation rates in Internet surveys. Barrios et al. (2011) document a higher participation rate in the web-based survey relative to a response rate in traditional mail-in surveys. However, this comparison was executed on samples of PhD graduates, who need not be representative of population at large in terms of time available, computer literacy, and so on. Indeed, it appears that the results so far have been mixed: There is no consistent evidence that web-based surveys differ on average in the propensity to participate from the traditional surveys (Fleming and Bowden 2009; Shannon and Bradshaw 2002; Shin, Johnson, and Rao 2012). 7
In addition to the response/participation rate, literature questions the randomness of the very decision to participate in the survey. Heiervang and Goodman (2011) claim that if the decision to participate in the survey is random and population is large enough, low response rates need not be a serious issue. They further argue that researchers should really be concerned about the quality of the responses and, consequently, the quality of the obtained results. On the one hand, access to Internet, computer literacy, and interest to participate in surveys are rarely evenly distributed in the population (see Chen 2014; Valliant and Dever 2011, for a recent overview of methodological approaches). On the other hand, some topics may particularly encourage/discourage certain types of responders to participate at all (Fang, Shao, and Lan 2009; Wright 2005) or to complete the survey (Tijdens 2014).
Another challenge concerns the quality of the data. Without surveillance while filling the survey, risks associated with satisficing may intensify (Stolte 1994). This problem arises when, fatigued by the survey, respondents provide answers that require less effort, for example, lining answers in a series of multiple choice questions, rounding up responses, and so on. This issue occurs in general self-administered surveys to a larger extent than in personal interviews (Sue and Ritter 2012). Revilla and Ochoa (2015) show that satisficing is associated with faster completion time of the surveys, which suggests that information on the length of the survey could be used as a proxy for the quality of the answers. As with the case of participation/response rate, research on the quality of responses provides mixed results. Some studies showed that web-based survey provided data of better quality, and, therefore, the results are more reliable (Braunsberger et al. 2007; Roster et al. 2004); others suggest that web-collected surveys generate more useless data (Cole 2005).
Representativeness of the Online Data
One of the issues that are most often raised by researchers is representativeness of the sample, a problem that is particularly acute in the case of volunteer surveys (Couper 2000). In spite of Internet’s increasing penetration, access to the Internet is still unequally distributed. In some countries, individuals with lower earnings might be underrepresented among Internet users. Similar arguments might be put forward in the case of elderly or less educated individuals. While in the developed world, these concerns might play a smaller role, in developing nations, they cannot be ignored (James 2008; Tijdens and Steinmetz 2016). Even in the cases where surveys reach entire targeted population, nonresponse might be larger in the case of Internet-based surveys due to technical issues, such as the lack of a stable connection or insuficient time to complete the survey, among others. Finally, one should consider that individuals of different groups might have different preferences concerning Internet use (e.g., Chen 2014; Steinmetz et al. 2013).
An alternative approach to explore the bias from online surveys relies on dual databases, that is, databases that contain two modules: one administered to a representative sample of the population and a web-based module. Bandilla, Bosnjak, and Altdorfer (2003) is an early example of this type of analysis, as authors compare two samples of the ISSP in Germany. They show that participants differ significantly with regard to sociodemographic structure and in other relevant variables. Even after weighting, results were inconsistent across samples. However, when they repeated the procedure within educational categories, differences in descriptive statistics disappear. By contrast, Schonlau et al. (2009) show that the use of propensity score matching weights improves the fit between the representative and the web-survey based components of the U.S. Health and Retirement Study, though some differences remain. Differences in distributions between samples from web-based surveys and traditional, representative surveys were also observed in the American National Election Study (Malhotra and Krosnick 2007) and in two similar surveys (one conducted via Internet, second as face-to-face interview) in Belgium (Loosveldt and Sonck 2008).
The Case of WI
Utilizing experience from the other online surveys, WI appears to be particularly concerned with the quality of the collected data and the outreach to the relevant population. Indeed, WI data are collected by experienced researchers and with great attention to methodological prudence; hence, their quality is possibly much better than ad hoc surveys in many countries as well as quasi-commercial data on wages from various wage comparison/ranking tools. WI data are collected through multilingual websites, one for each participating country. In some countries, WI also provides websites targeted toward groups such as women, specific professional groups, and so on. 8 Websites that offer WI survey have to fulfill a “quality of information” criteria set by the nongovernmental organization operating WI. These websites rank high in search engines for a wide array of key words. Hence, WI recruitment is based mainly on voluntary participation by individuals suficiently interested in related topics to query one of the related key words in the search engine. In addition, many specialized service providers—for example, job brokers, temporary work agencies—advertise the tools offered by WI.
Recognizing the risk of satisficing (see Sue and Reiter 2012), WI survey is short and only few questions are actually obligatory to the participants. Moreover, participants are incentivized to complete the survey, both because they can get a more accurate SalaryCheck data and they get a chance to obtain a monetary prize equal to the weekly minimum wage, monthly in the case of countries with low minimum wages. Chances are doubled for participants willing to become part of a panel survey. 9 The standard version of the survey requires approximately 10–20 min to complete, but in countries with slower average speed of Internet, this survey is further shortened to roughly 5 min (Tijdens et al. 2010). Finally, the survey can be completed in several sessions within a week span.
Given the extensiveness of the WI data in terms of breadth and coverage, earlier literature on the data project asks whether WI data are reliable enough to be used for research purposes. In analyses for Germany and the Netherlands, Steinmetz et al. (2009) and Steinmetz and Tijdens (2009) compare the WI data to nationally representative databases, the German Socioeconomic Panel and the Labor Supply Panel from the Netherlands. They find that in both cases, WI samples are not representative of the general population. Consequently, wage distributions differ between the two types of the sources, yielding biased estimators in the case of online sources vis-à-vis the representative sources.
In summary, previous findings consistently report differences between the representative and the web-collected surveys. The use of weights to improve the fit between the samples presents a mixed record in terms of balancing the data from online surveys. The main limitations of the previous literature concern both methodology and the coverage. In terms of methodology, the construction of weights does not assure balancing and necessitates access to both WI and benchmark data, as both WI and benchmark samples are reweighted. In terms of coverage, the reliability analyses are available for selected countries and years. Against this background from the earlier literature, our analysis contributes to both the methodological and the cross-country dimensions of the WI. We test a novel procedure for balancing WI data to nationally representative benchmark samples. This new approach assures balancing and reweights only the WI sample, which makes it a convenient addition to the raw WI data for all researchers interested in making a relevant WI sample resemble a representative population despite lack of access to a representative sample. Wide country coverage allows researchers to explore more systematically the reasons behind WI representativeness or lack thereof. While in a strict sense, we can only comment on the actual 95 data sets that we use in the analysis, the broad coverage of countries, years, and various sources of data lends some grounds to cautious generalizations of our results.
Method
This section discusses the methods employed in our study. We first briefly describe the statistical tests used for the comparison of the WI data and a benchmark data set. In a second part, we review the reweighting procedure employed in our analysis.
Consider a benchmark sample from a population that is representative along the defined criteria of representativeness. Typically, for nationally representative samples, residence, age, and gender are considered suficient criteria for random sampling from the population by most central statistical offices around the world. Such approach hinges crucially on the implicit assumption that conditional on matching these characteristics between the random sample and the population, the measurement of other characteristics is as good as if each individual from the population participated in the survey (notably with a sampling error declining with the number of participants). Often, nationally representative surveys utilize administrative records to know the “true” geography, age, and gender distributions of individuals and subsequently randomly sample addresses to perform the questionnaire. The random sampling is key to assuring that each individual within society has the same probability of participating in a survey. Stratification is used to mitigate the risk that the sample population is excessively dominated by this strata of the society that is the easiest to access. Since participation in the questionnaire is never fully warranted, the realized distribution of the key characteristics is used to obtain weights which make the sample representative of the underlying population. If nonparticipation is random, weights are neutral. If nonparticipation is larger among specific strata of the society, survey weights correct for that fact.
Against this benchmark case, consider an alternative sample, for which the sampling procedure is unknown, but the final distribution of the key characteristics is known. Such surveys are sometimes referred to as nonprobability surveys. If one is able to provide a structure of weights that makes this sample from an unknown sample design replicate the distribution of the individual characteristics from the representative data, one can extend the argument from the nationally representative sampling: conditional on these weights, answers to all other questions should be as good as if each individual of the population was asked, with a sampling error. Of course, this is only warranted if the (unknown) sampling design is not affecting the measured characteristics themselves. After reweighting, answers should be a correct approximation of the underlying population, conditional on participation being independent of a given analyzed characteristics. 10 Naturally, the sampling error cannot be obtained if the sampling design is unknown.
In this article, we compare samples from the WI (which fits the description of the alternative sample) to benchmark samples (as discussed above) in order to obtain the weights which help to correct for the unknown and thus possibly nonrandom sampling in the WI. To this end, we collected a large number of nationally representative samples in terms of key characteristics: age, gender, and residence. These samples are subsequently compared to the WI, conditional on the underlying characteristics. Since both the nationally representative surveys and WI comprise a large number of outcome variables in addition to the population characteristics, we select one such variable—hourly wage—to analyze to what extent WI can indeed be comparable to nationally representative sources.
Comparing the Distributions
The principal interest lies in testing if data from two sources come from the same distribution. This analysis poses two important challenges. First, the sample sizes in the two sources may differ substantially. Second, self-reported data, such as WI or LFS, are likely to contain more round numbers, whereas administrative sources, such as structure of earnings survey, are likely to contain exact gross wages, which are rarely round. These two challenges necessitate that the tests to be employed make no assumptions about the underlying distribution of the data. Such requirement yields three candidate tests to compare the two samples: Kolmogorov–Smirnov test, Mann–Whitney U test, and Epps–Singleton two-sample test. These three tests share the null hypothesis that both analyzed samples are drawn from the same population. However, the power of the tests varies.
The Kolmogorov–Smirnov test (Kolmogorov 1933; Smirnov 1933) is the most widely used among the three, and it is based on a comparison of the cumulative distribution function in the two samples. The test statistic is proportional to the maximum discrepancy between the two samples. This test is sensitive to differences in the median, the shape, and the span of the distributions. It permits the use of survey/sample weights. However, its properties rely heavily on the assumptions of continuity of the distributions (the distributions should be fully specified). Moreover, it tends to be less sensitive if the discrepancy occurs in the upper tail of the distribution.
The Mann–Whitney U test (Mann and Whitney 1947) follows a different approach. Instead of comparing cumulative distribution functions, it ranks all observations. Under the null hypothesis, that the two samples came from the same population, the ranks will be randomly distributed between the two samples. This implies that this test is better to capture changes in the location of the distribution, which are usually reflected at the median. In addition, Schmid and Trede (1995) demonstrated in a Monte Carlo experiment that if wages follow a Pareto distribution, the Mann–Whitney U test is more powerful than Kolmogorov–Smirnov. 11
Finally, the Epps–Singleton (1986) test is based on the empirical characteristic function. When compared to the Mann–Whitney U test, the main advantage of Epps–Singleton lies on its ability to detect discrepancies in the location, family, and scale (Goerg, Kaiser, and Bundesbank 2009). When compared to Kolmogorov–Smirnov test, Epps–Singleton has two advantages. First, it is more flexible since the characteristic function is completely defined for discrete and continuous data. Second, it tends to be more powerful (Goerg et al. 2009).
Given the advantages and disadvantages for each test, we employ all three, adapting them to accommodate for sample weights. We provide systematic tests for each analyzed sample, with the null hypothesis that both WI sample and the alternative sample are drawn from the same population. We provide these tests for the raw WI data as well as for the WI data after applying the reweighting procedure as described below.
Reweighting Procedure
If tests reject the null hypothesis that data come from the same underlying distribution as the population, then suitability of the nonrandom sample from the population for reliable statistical inference without any further correction becomes questionable (Valliant and Dever 2011). A possible solution to this problem is to reweight observations in the WI to make them resemble representative data. Steinmetz et al. (2009) use (the inverse of) the estimated propensity scores as weights.
12
Formally, the procedure involves running a probit/logit regression where the explained variable is the source (a binary variable that takes the value of 1 when an observation comes from benchmark data and 0 otherwise). Then, one may define
Moreover, the estimation of the propensity score with the use of the probit/logit maximum likelihood estimation is in fact inferior to alternative methods, in particular, in the case of nonrandomly missing data (e.g., simulation and data examples from Imai and Ratkovic 2014). Hence, a better solution is to rely on moment-based estimation in obtaining the propensity score. This approach was proposed by Imai and Ratkovic (2014) and yields propensity scores that, by construction, balance the covariates, hence the name covariate balancing propensity score (CBPS) matching. The procedure is immune to the propensity score misspecification problem, as it exploits the dual nature of the propensity score as a covariate balancing score and the conditional probability of assignment to subsample. 14 Imai and Ratković (2014) notice that estimating the propensity score via maximum likelihood, as used typically, can be conveniently rewritten as a transformation of the sample moment conditions for the covariates that are used to obtain the propensity scores. In other words, the propensity score can be thought of as the (non)linear combination of individual characteristics that maximizes the probability that observations are correctly assigned to a subsample. Hence, one can recover weights that produce an exact balance of covariates, in our case—between the WI and the benchmark representative samples.
Conveniently, the derivation of the propensity score via moment-based estimation also gives clear interpretation for the theoretically warranted specification of the weights. Recall, that earlier studies used IPS, which by definition cannot balance samples from benchmark and WI data. In contrast, we rely on the theoretical result of Imai and Ratkovic (2014) and provide the weighting scheme which adjusts data from WI to reflect the structure of the sampling design used in when obtaining the benchmark sample. Specifically, the weights imposed on benchmark representative samples are equal to 1 for all observations in this data (conditional on utilizing survey weights in estimating the propensity score). Adapting Imai and Ratkovic (2014), this necessitates the following weights for the WI data:
where N denotes total number of observations from WI and benchmark representative data, NB denotes the number of observations in the benchmark representative data (B), and PB (Xi ) indicates the score for an observation i, obtained using the moment based approach offered by Imai and Ratkovic (2014). This score is analogous to the probability that observation i was obtained from the benchmark sample, given its characteristics.
A conventional alternative to matching using CBPS is the maximum likelihood estimation of propensity scores with subsequent use of balancing weights. Given typically large sample sizes of the benchmark representative data relative to WI data sets, kernel weights appear as superior. 15 With kernel density (KD) matching, the weights for particular observations represent the distance between its propensity score and the scores of the observations from the benchmark sample. Formally, we follow Smith and Todd (2005) and Morgan and Harding (2006) and calculate the weights as:
where
KD weights display two main advantages relevant to our context. First, they do not require the researcher to make any arbitrary restrictions on how many and which observations to select from the control group. In fact, the computation of weights happens automatically for all observations. This leads to the second advantage of using the KD weights: A computed weight is a synthetic measure for each observation from WI of how similar it is to all observations from a representative sample. Thus, risk associated with bad matches is minimized, while each observation from WI may be included in subsequent analyses.
Clearly, one would want to balance WI and the benchmark representative data on the same variables, as are part of the sampling design for the benchmark representative samples: place of residence, age, and gender. However, information on place of residence is often missing in WI or is reported in a way that does not permit straightforward comparison with the benchmark representative data sets. 16 Also, our interest in this article lies in salaries. Hence, we decide to include education in addition to age and gender in the matching procedure. While other human capital variables are asked for in the WI surveys—such as tenure, experience, occupation, or industry—the proportion of missing values for these variables is much larger, a problem shared with many benchmark samples. Their inclusion as additional covariates would have led to a significant reduction in the sample size from the WI, and therefore we confined the matching variables to the most widely accessible. Summarizing, we use age, gender, and education to obtain both weights: conventional propensity score with kernel matching and CBPS.
Testing for the Bias After Reweighting
We perform an Oaxaca–Blinder-type decomposition as operationalized by Jann (2008) 17 on a Mincerian wage regression. We decompose the difference between WI and each benchmark sample into a part that is attributable to differences in individual endowments between the two data sets, also known as “explained’’ component, and a part that remains attributable to differences in the coeficients when wage regressions are estimated separately for each data set. This second term is the “unexplained’’ component. 18 Since this decomposition is based on a regression approach, the use of weights is nonproblematic. 19
Performing an Oaxaca–Blinder decomposition has two main advantages. First, it provides an additional test of the quality of the balancing: A successful balancing implies that differences in characteristics should vanish in statistical terms. This is equivalent to stating that the explained component of the difference between wages from benchmark data and WI should be negligible. Hence, all the difference should be related to the unexplained component, that is, to differences in coefficients. Second, the Oaxaca–Blinder decomposition allows distinguishing the contribution of each of the covariate to total differences in the coeficients, thus making it possible to identify the sources of the eventual differences between WI and the benchmark samples. 20
An issue that arises with the use of Oaxaca–Blinder decomposition is the choice of the structure of wages to be considered as counterfactual for testing the hypotheses on the obtained parameters. We use the parameters from the benchmark samples. This choice is motivated by the key research question behind our article. Thus, the counterfactual represents the wage that participants of the WI would have recorded in the benchmark samples if their characteristics were valued according to the same schedule as in benchmark samples. 21 Hence, a decomposition of the following form is run on reweighted data:
where X denotes the individual characteristics (i.e., age, gender, and education). In this approach,
Data
The WI project pioneered large-scale wage data collection directly from online questionnaires. The first results were already available in 2000. Initially, the project was restricted to the Netherlands, as the survey was coordinated by the University of Amsterdam. In 2005, eight European countries joined the project. Since then, the number of participating countries increased to reach 96 countries from all over the world. 22
In many regard, the questionnaires provided by the WI survey resemble those employed in traditional surveys, particularly LFS. Respondents provide answers on wages and a large number of individual characteristics such as year of birth, gender, occupation, household characteristics, and so on. It also covers topics that are usually not included in standardized surveys, such as characteristics of the current employment, workers’ attitudes and satisfaction, overeducation, and so on.
Nationally representative surveys are typically dificult to obtain and country-specific. We benefit from a large collection of harmonized nationally representative data sets such as LFS and HBS. In most countries where LFS and HBS are available, they come from random sampling from the population based on age, gender, and residence. There are also alternative data, whose representativeness is warranted within the population used for sampling. An example is European Union Structure of Earnings Survey (EUSES), which comprises salaried workers within a segment of the enterprise sector defined independently for every country. The most frequent sample design comprises complete coverage in firms employing between 9 and 49 workers and random sampling within firms employing more than 49 workers, full-time equivalent. EUSES data do not cover salaried workers from public institutions, neither in elected nor administrative positions. Weight design in EUSES allows researchers to generalize the surveyed population to all employed workers in the private sector. Hence, EUSES is not representative of the entire population. 23 Finally, we also employ large-scale random sampling surveys, following a coherent methodology. An example of such survey, which collects information partly analogous to WI, is ISSP. While ISSP typically has smaller sample size than LFS or HBS, individuals are randomly sampled from the population.
The representative data sets that we use as a benchmark come from all three types of the abovementioned sources. First, we use the linked employer–employee data of administrative quality about gross wages. This type of data is available for Hungary from the country’s Central Statistical Office as of 1992 (structure of earnings survey [SES]) as well as for all members of the European Union available from Eurostat as of 2002 (EUSES). Second, we use LFS and HBS collected by the central statistical ofices of Argentina, Belarus, France, Germany, Great Britain, and Poland. These are self-reported wage data, with large and nationally representative samples. Finally, we also employ self-reported data from the Russia Longitudinal Monitoring Survey (RLMS), the German Socio-Economic Panel (GSOEP), the British Household Panel Survey (BHPS), and the ISSP. While samples from the latter source are often smaller relative to LFS or SES, nationally representative sampling was used to collect the surveys.
Such a large selection of microlevel data sets permits a comprehensive comparison of WI data to benchmark representative samples. For a given year in WI sample, we rely on more than one representative database, with differentiated sample size and designs. For example, SES is administered to employees of the enterprise sector, in some countries with the additional restriction that the employer has to be characterized by a suficiently large number of employees. We utilize the same restrictions when matching WI data to these sources. In those cases, if needed information is missing—for example, WI has no information on the industry of employer—the observations are dropped from the WI sample.
For the comparisons to be meaningful, we utilize WI samples that have a suficiently large number of observations to maintain statistical properties. We set the threshold to at least 100 complete records in WI, that is, complete information on age, education, and gender. 24 Within a large collection of the individual-level microdata sets, we select those which match to WI in terms of country and year. Tables A1–A4 in the Online Appendix report the detailed list of sources and years for each analyzed country. In total, we obtain 95 matching year and country representative data sets from 17 countries (92 for hourly wages). This collection of surveys represented in our study include advanced, catching up and developed economies from Europe, both Americas, and Africa. 25 We describe the benchmark data sources in more detail in Online Appendix A1.
In both WI and the benchmark data, the wages are reported in local currency unit from the same period, which makes comparisons immune to issues such as currency conversion or inflation. Wages are typically reported in weekly, monthly, or hourly manner. If only monthly wages were reported in the benchmark sample, we convert them to hourly wages by dividing monthly wages with weekly reported hours of work times 4.33. Similarly, if only weekly data are available in the benchmark sample, we convert it to hourly wages by dividing the weekly rate by the reported number of hours. In the case of three data sets, wages are reported in monthly terms and no data on hours worked are reported. These three data sets are dropped from the analyses, but for comparison purposes and as a robustness check, we also repeat the tests for monthly rather than hourly wages.
In parallel to wages, age and education measures also were harmonized between WI and the representative benchmark data. Age variable was recoded to age groups, commonly defined in all data sets. For education, we harmonize the information about educational attainment to three classes: tertiary or above, secondary, and primary or below. We consider vocational education to be secondary education. Since we only match WI data to the data from the same country and the same year, country- and time-specific features concerning, for example, the role of vocational, secondary, or tertiary education do not affect the quality of the matching.
Results
First, we show the outcomes of tests for the equality of wage distributions. These analyses are performed before balancing the samples. We then show the results of balancing and subsequently move to analyzing the differences in the schedules of wages after balancing. We compare the samples on the basis of two main outcome variables: hourly wages and monthly wages. When available, we use the actual indicator of hourly wages from the survey (WI or nationally representative).
Differences in Wages Before Reweighting
Wage distributions from WI are in a vast majority of cases different from the distributions in the nationally representative data as documented in Table 1. One obvious way to compare the two data sets is a simple statistic for the means from the two distributions to be equal. These tests show that 11 samples of 95 have statistically similar means. However, such tests are unable to capture differences at other points of the wage distributions. We proceed to complement them with the tests described in Method section. These results confirm that (hourly) wages in WI and nationally representative data differ in statistical sense. In fact, we reject the null hypothesis of wages coming from the same distribution in more than 95 percent of the cases. Rejection rate is actually within what it was expected for significance test with 5 percent confidence level.
Tests for Equality of the Wage Distributions for WI and Benchmark Data.
Note: The table presents predicted differences in (log) median hourly wages between Wage Indicator (WI) and nationally representative samples. Predicted values come from a regression where the dependent variables is the difference in average hourly wages as a percentage of hourly wages in the representative sample. Values above 1 indicate higher wages being reported in WI. Regression also includes controls for country and source. Regression does not control for differences in characteristics between respondents of WI and nationally representative samples. H0 = null hypothesis.
With WI becoming more recognized and more reliable, one could expect that the rejections of the null hypothesis become more unlikely. To test explicitly this hypothesis, we estimate a model with the mean difference between WI and benchmark data as an explained variable. 26
The set of explanatory variables contains nothing but fixed effects for country, data source, and year. Hence, we may obtain conditional predictions of the difference for the consecutive years covered by WI that are “clean” of the country specificity and data source specificity. The results are reported in Figure 1 in the form of the conditional predictions with confidence intervals of 95 percent. Points below the zero line correspond to mean hourly wages in WI short of analogous value in benchmark nationally representative data. Time trends display no specific pattern. In fact, the differences in mean hourly wages tend to be large at all times, despite substantial increase in WI sample sizes.

Hourly wages: Distribution of differences between Wage Indicator (WI) and benchmark data. Figure presents predicted differences in (log) median hourly wages between WI and nationally representative samples. Predicted values come from a regression where the dependent variables is the difference in average hourly wages as a percentage of hourly wages in the representative sample. Values above 1 indicate higher wages being reported in WI. Regression also includes controls for country and source. Regression does not control for differences in characteristics between respondents of WI and nationally representative samples.
As noted at the beginning, differences in wages could be a reflection of differences in sample composition. As suggested in the literature, differences in Internet access coupled with preferences of the individuals concerning its use could result in WI samples characterized by relatively younger and better educated individuals. To identify sources of differences, we estimate the mean values of several characteristics of interest (age, education, and gender) on country, source, and year fixed effects. Fitted values for the latter are plotted in Figure A1 in the Online Appendix. 27 In the case of age, differences appear to be widening over time: recent waves of WI are on average 5–10 years younger than those in the representative sample. Similarly, we observe that participants in the WI are on average better educated, as the proportion of respondents with only primary studies is smaller in virtually all cases. Since the selectivity patterns appear to be systematic, we move now to balancing the WI samples to resemble nationally representative distributions in terms of human capital characteristics: age, gender, and education.
Balancing WI Data
We employ three key human capital indicators: gender, age, and education. We make sure that the sample design of the benchmark nationally representative data is reflected in which observations from WI are used. For example, if samples of SES and EUSES cover private employers with nine or more employees, full-time equivalent, we exclude individuals who do not meet these criteria from the WI data prior to matching. Hence, in those cases, we work with a subpopulation of WI rather than the complete data set.
We balance the distributions using the two approaches discussed earlier: Imai and Ratkovic (2014) estimator and KD matching estimator. To facilitate comparison, we provide tests also for the raw (unweighted) distributions and for the weighting scheme suggested by Steinmetz et al. (2009). The results are reported in Table 2, portraying the summary of variable-by-variable, pair-by-pair testing of balancing. 28 The results reveal that, in principle, WI and nationally representative data differ substantially, which was hinted already by Figure A1 in the Online Appendix. Then, the method proposed by Steinmetz et al. works to some extent with the ISSP data but in some cases may actually reduce the balancing. Weights derived from KD matching on a propensity score similar to Steinmetz et al. do better for the ISSP data, balancing majority of these samples. Admittedly, it is not as effective for other sources of data. Finally, our preferred weighting scheme, based on CBPS, is able to balance all the sources of the data. This result is embedded in the estimation method and thus should come as no surprise; but, in the context of the other methods, it shows the improvement in balancing that may be achieved by changing how weights are computed.
Balancing of the Characteristics Between Wage Indicator and Benchmark Nationally Representative Data.
Note: The table shows the frequency in the rejection of the null hypothesis that sample is balanced. None signifies the number of samples where all covariates are balanced, all signifies the cases where no covariates are balanced, and 1–2 signifies the number of samples where some covariates are balanced. Column titled Kernel Density Weights obtains weights from propensity score using the kernel matching algorithm. Columns titled IPS Weights use inverse propensity score weights, which replicate the approach employed in Steinmetz, Tijdens, and de Pedraza (2009). In both cases, propensity scores were obtained from a probit regression on age, gender, and education level. Weights were obtained for all databases, including those for which later analysis of the wage structure is not performed, for example, where benchmark data contain only categorical information on wages. Hence, we report results for 123 data sets at hand, whereas the remaining analysis is performed for 92/95 data sets, for which continuous information on wages is available (see Tables A1–A4 in the Online Appendix for details on data availability per country, year, and data source). CBPS = covariate balancing propensity score; RLMS =Russia Longitudinal Monitoring Survey; GSOEP = German Socio-Economic Panel; BHPS = British Household Panel Survey; ISSP = International Social Survey Program; EUSES = European Union Structure of Earnings Survey; IPS = inverse propensity score.
While the use of weights improves the balance of characteristics across samples, results for wage distributions are less encouraging. Repeating the exercise from Table 1 reveals that weighting with our preferred weights has some small effect on the match between the distributions of wages, see Table 3. In fact, there were three cases for wages and five cases for hourly wages when WI distributions were found to match the nationally representative data for the Mann–Whitney test and individual such cases for the other tests.
Wage Distribution After Weighting With CBPS Weights.
Note: The null hypothesis (H0) states that both samples were drawn from the same distribution. The alternative hypothesis (H1) indicates rejections of the null hypothesis at the 5 percent confidence level.
In order to better understand to what extent the remaining differences in hourly wages are related to different wage schedules in WI relative to nationally representative data, we proceed to perform the Oaxaca–Blinder decompositions. We provide two alternative specifications for the unexplained component: with and without a constant. Such a choice is motivated by the fact that WI has two measures of wages: gross and net. By contrast, nationally representative data sets usually contain only one measure, either gross or net. What is more, countries differ in what is exactly the difference between the gross and the net. 29 Finally, in some of the countries in the sample, it is customary to contract on net wage (tax and social security contributions are effectively paid by the employer), whereas in others, it is the gross wage that is more socially embedded. If the difference in the distributions between WI and nationally representative data was somehow driven by the confusion between gross and net wages, the specification of the Oaxaca–Blinder decomposition without a constant is able to accommodate for this fact. Admittedly, the differences in the constant might come from various sources, for example, differences in the survey design (specific phrasing of the question about wage), preference toward rounding earnings figures, and so on. They all can display as differences in the constant between the two data sources. We keep the same human capital variables that we employed in the propensity score matching: age, education, and gender.
The results reported in Table 4 reveal that excluding a constant from the unexplained component of the Oaxaca–Blinder allows to achieve as many as 36 unbiased pairs of samples (of 92) for hourly wages, of which 28 were obtained for balanced covariates and 8 despite the lack of balancing in the covariates. There are three data sets more if we analyze the conditional wage distributions instead of hourly wages. Note that the balancing weights are obtained for all the salaried workers, whereas the estimations of Oaxaca–Blinder (similar to the distribution tests discussed above) are only possible for salaried workers who report wages. The problem of missing information on wages is more pronounced in the survey benchmark representative data, hence making the sample participating in the regression different from the sample for which the balancing is obtained. By contrast, WI data typically always contain information on wages. This hints, that depending on the objective, researcher may want to balance the WI data to general characteristics of the analogous population in the representative data or to the population with similar information coverage, especially on wages. Notably, the two need not perfectly overlap. The more selective the information on wages in the nationally representative data is, the less similar the samples to the WI data, regardless of the differences in the survey design and data collection.
Oaxaca–Blinder Decompositions After Weighting.
Note: The H0 states that the joint estimate of the differences between samples in a pair is statistically insignificant (at the 5 percent level). Rejection of this hypothesis (H1) states that endowments and/or coeficients differ between the samples in a pair. Specification with a constant includes constant from Mincerian wage regression in the test for equality of coeficient as part of the unexplained component in the Oaxaca–Blinder decomposition. The opposite holds for a specification without a constant. CBPS, Kernel density, and IPS indicate three weighting schemes used to balance covariates. Please see Table 2 Note for more details. H0 = null hypothesis; H1 = alternative hypothesis; CBPS = covariate balancing propensity score; IPS = inverse propensity score.
Comparing CPBS weights to the alternative schemes reveals that CBPS outperforms all others. In comparison to kernel weights, CBPS is able to make three additional databases conditionally similar. The difference is much larger for the IPS method proposed by Steinmetz et al. (2009): roughly 19 or 20 samples more are made conditionally similar (in the case of hourly wages and wages, respectively). This means that roughly half of samples cannot be effectively reweighed in terms of characteristics using the IPS method but can be effectively reweighed with CBPS weights.
The number of unbiased estimators goes down to as few as 20 if we allow the constant to be a part of the unexplained component. These results suggest that differences in wage levels between the two samples (reflected in the constant) are important drivers of the unexplained component, while the marginal effects of human capital variables in the reweighted WI and the nationally representative data appear to be fairly comparable. Notwithstanding, significant differences occur more frequently in the comparisons to EU-SES and national LFS. In Tables A1–A4 in the Online Appendix, we present detailed results for the different types of data sources. 30
Overall, our results suggest that WI data should be used with caution, even after reweighting. One of the reasons for failure to reject the null hypothesis that the estimates from nationally representative data sets and WI are the same can stem from a relatively lower precision of the estimates obtained for WI. Arguably, with a smaller sample size, the estimates of the coeficients are likely to have wider confidence intervals. For very large sample size in WI, even if statistically significant, the differences are economically small. The opposite tends to hold for small sample sizes. Another interesting insight is that wages in some countries tend to be overstated in WI. Moreover, these deviations appear as large, 50–80 percent of mean/median wage. This pattern might reflect a self-selection process: People whose wages are systematically high for unobservable reasons, or at least those who report such wages, appear to be more willing to participate in the WI survey in some countries. Increasing the popularity of WI data is likely to reduce its selectivity both in terms of observable and in terms of unobservable characteristics. It is worth to mention that WI project was originally focused on specific labor market problems connected with discrimination, for example, gender wage gap. Respondents are encouraged to take part in the survey in order to compare their salaries with similar respondents and check if they should earn more (e.g., SalaryCheck and minimum wage tools of WI referred to in Insights From the Literature section). This feature of WI should have attracted those who expect that they might earn less.
This selection on unobservables appears to be particularly strong for several countries, for which the difference does not disappear even once the weighting is implemented. Namely, in Table A5 in the Online Appendix, we report the estimated effects of a given country, controlling for year and data source. The explained variable in this regression is the difference in wages between WI and benchmark representative data (expressed as a percentage of the mean in the representative data), by analogy to results reported in Figure 1, however, with two main differences. First, we include the regressions for the reweighted distributions. Second, we also show the results at the mean (to be analogous to the results from Oaxaca–Blinder results). Notably, some countries have significant fixed effects even after reweighting—for example, Australia, Germany, or Italy—whereas for some others, the reweighting makes the difference statistically close to 0—for example, Russia or Ukraine. Finally, for some selected countries, the differences did not appear to be statistically different from 0 even prior to reweighting—for example, the Netherlands, Sweden, or Finland. However, given that our sample includes 17 of 90+ countries covered by WI, it would not be grounded to form judgment concerning the specificity of some countries.
To identify the patterns that could stand behind the country specificity and thus explain the scope of difference between WI and the benchmark nationally representative data, we run a toy analysis, where we regress the differential in log median wages (raw and reweighted with CBPS weights) on two country characteristics: income (proxied by gross domestic product per capita) and Internet use (proxied by Internet penetration statistics). These results are reported in Table A6 in the Online Appendix revealing that higher income countries with more widespread Internet use tend to be characterized by lower differentials. These results might come either from a better matching or from an increase in the representativeness of the sample. In either case, it is possible to be optimistic about the future of the WI. The differences between WI and nationally representative data should shrink over time, for example, due to the increase in the Internet penetration or to the higher awareness of the existence of the WI project. Both trends might increase participation in the WI and allow more analyses using these data. However, after correcting for differences in age, gender, and education (i.e., after reweighting), neither variation in Internet penetration nor the income per capita has the explanatory power in the regressions. This suggests that while higher income and widespread Internet access may make participation more universal, WI data still have other selectivity patterns—unrelated to those observable characteristics—that drive systematical wage differentials.
Conclusion
Internet offers great opportunities for researchers to gather dedicated data, but it also poses potential dificulties. A critique of online surveys focuses on sampling: data from such sources do not have to be representative of the underlying population. Consequently, statistical inference and external validity are sometimes put in question. This problem might be particularly acute for social research as online surveys are often the only possible source of data. Administrative data or nationally representative public surveys hardly ever include questions on life or job satisfaction, work–life balance, feelings or attitudes, and so on, which seem to be crucial for studying life conditions and life quality, whereas executing a dedicated nationally representative data is often prohibitively expensive. With the growing popularity of Internet and growing sample sizes in online surveys, many argue that the problem of representativeness is becoming less relevant. In this article, we provide empirical evidence of this conjecture. Specifically, we investigate the reliability of data from the WI—a large-scale multinational online survey. This survey covers a wide range of countries for a relatively long time span and contains a comprehensive set of variables including human capital variables, employment characteristics, and satisfaction with different aspects of the job. Given its richness, the WI has enormous potential for research on labor for sociologists, psychologists, economists, and anthropologists alike.
In order to assess its reliability, we compare data from WI to 95 nationally representative surveys from 17 countries—both industrialized and developing. We analyze the wage distributions and the individual characteristics. The results of this comparison suggest that participants of the WI do not come from a representative subpopulation. This different sample composition translates to differences in the distribution of wages and hourly wages. The key contribution of our study is to provide a novel method to reliably ameliorate the differences in the sample compositions between WI and benchmark representative samples for these countries via reweighing. Our method draws on the recent developments in statistics, namely, a CBPS estimator. The provided weights reduce the discrepancies in the individual characteristics across WI and benchmark nationally representative samples.
However, despite balanced populations, the reweighted wage distributions continue to differ. In fact, on average, WI respondents tend to report higher wages than conationals with similar characteristics in representative surveys. Namely, WI respondents tend to be younger and better educated than a representative sample. Yet, compared to identical individuals from representative samples, WI respondents tend to report higher earnings. This feature holds for a large share of analyzed countries and years. It is beyond the scope of our study to determine whether this disparity stems from systematic overreporting in WI or self-selection into participating in WI. Hence, despite successful rebalancing of the WI samples’ structure, our results cast a shadow of doubt on the use of WI data to obtain estimates with reference to the entire population of the countries participating in the WI. On the positive side, the proposed weights help to bring WI closer to the nationally representative samples in a large number of cases, as the estimates of the Mincerian wage regression from WI cease to be biased relative to the representative samples for many of the countries covered in this study. On the negative side, we find no confirmation that WI becomes closer to benchmark nationally representative data over time per se. One would typically expect that once more users become used to this form of surveying and the more common it becomes, the more similar WI data should be to traditionally administered representative surveys. It appears that selectivity patterns associated with income and Internet access can be meaningfully corrected with the proposed weighting scheme. What cannot be corrected is the sample selection on unobservable characteristics: individual characteristics that affect both wages (and potentially other answers in WI) and the very participation in the survey.
The key caveat is that no weighting procedure to balance one data set to replicate the structure of another data set can make up for the observations missing in either of the sets. Although trivial, this fact is of paramount importance for assessing the applicability of online survey data for social research. While the balancing properties can be satisfied to make WI data resemble the nationally representative data for the observables—reweighting will leave three important issues unaddressed. First, self-reported data sources such as LFS, HBS, and many social studies suffer from incomplete coverage on specific questions. This incomplete coverage may occur systematically, which confronts the researcher with the decision what is his benchmark sample for balancing the online survey data. This choice may have nontrivial consequences for the results. Second, online surveys may attract participants selectively on characteristics unobservable in the nationally representative data, but which are relevant for a given variable of interest in social or economic modeling. If that is the case, the obtained estimates remain biased even after weighting. Third, for some online survey data, the existing counterparts come from populations which purposefully do not fully overlap in terms of individual characteristics. If domains of the individual characteristics do not overlap, reweighting can only help to balance the matching subsamples from the two sources. Admittedly, these three issues—nonrandom selection in nationally representative data, nonrandom unobservable selection into online surveys and common domain—are of substance to many research projects and deserve further analysis. The procedure proposed in our study—matching on covariate-balancing propensity score—is able to address common domain selection on observables, leaving the researcher with flexibility on whether to balance to full population or a subpopulation of interest. The search for proper methods to address the remaining issues can improve further the reliability of the online survey data in providing insights into policy and social sciences in the future.
Supplemental Material
Supplemental Material, appendices - A Cautionary Note on the Reliability of the Online Survey Data: The Case of Wage Indicator
Supplemental Material, appendices for A Cautionary Note on the Reliability of the Online Survey Data: The Case of Wage Indicator by Magdalena Smyk, Joanna Tyrowicz and Lucas van der Velde in Sociological Methods & Research
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Lucas van der Velde gratefully acknowledges the support from the National Center for Science, project number UMO-2016/21/N/HS4/02108. Magdalena Smyk gratefully acknowledges the support from the National Center for Science, project number UMO-2016/21/N/HS4/02109 and the Foundation for Polish Science through the grant START 2017.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
