Abstract
Case studies reporting real-world experiences with survey falsification are uncommon. In this article, we document the experience of a panel survey in India that produced TV viewing estimates (“TV ratings”) where external parties were illegitimately trying to influence respondents’ behavior. The usual method to detect possible falsifications was that of analysts poring through data to find suspicious viewing patterns. Here, we develop a method using multilevel models and illustrate its use in the detection of an actual incident. We report how the model-based method was used to direct on-ground investigations that ultimately supported our analytic inferences. The model-based method offers four advantages over the usual method. First, by approximating an interpenetrated sample, the model simultaneously controls for several household characteristics. Second, Empirical Best Linear Unbiased Predictors (EBLUPs) of random effects can be examined separately at both the household and interviewer level, thus suggesting where further investigation efforts should be directed. Third, the method is faster and more objective than the usual method. Fourth, the method is easily implemented and can provide regular quality control for survey organizations.
Introduction
It is generally acknowledged that there is little literature on survey falsification (Bredl, Winker, & Kötschau, 2012; Koczela, Furlong, McCarthy, & Mushtaq, 2015) despite its far-reaching consequences (Office of Research Integrity, 2002). However, the last few years have seen a renewed interest in this area as witnessed by recent publications (e.g., Koczela & Scheuren, 2016; Winker, Kruse, Menold, & Landrock, 2015), conferences focused on this topic (e.g., National Opinion Research Center, 2016; New England chapter of the American Association for Public Opinion Research, 2015; Washington Statistical Society, 2014), and a special task force on falsification constituted by the American Association of Public Opinion Research (AAPOR) and the American Statistical Association (ASA). Activity has been spurred by development of methods to detect falsification (Bredl et al., 2012; Bushery, Reichert, Albright, & Rossiter, 1999; De Haas & Winker, 2014; Judge & Schechter, 2009; Murphy, Eyerman, Mccue, Hottinger, & Kennet, 2005; Porras & English, 2004; Swanson, Cho, & Eltinge, 2003); debates surrounding the actual prevalence of falsification (Kuriakose & Robbins, 2016; Simmons, Mercer, Schwarzer, & Kennedy, 2016); a realization that even if the prevalence of falsification is low, its impact on estimates can be large (Schraepler & Wagner, 2005; Winker et al., 2015); and looking beyond the interviewer to situations where, for example, survey staff use computers to fabricate data (Koczela et al., 2015).
Ideally, interviewers in a survey should be exchangeable, that is, holding all other survey conditions constant, we should obtain the same response from a respondent irrespective of which interviewer undertakes the interview. But in practice, answers from respondents interviewed by a particular interviewer tend to be similar to each other as compared with responses obtained from respondents belonging to other interviewers, even after factoring differences in geography and respondent profile (Biemer & Stokes, 1985; Fellegi, 1974; Hansen, Hurwitz, & Bershad, 1960). These measurement errors are called “interviewer effects.”
Multilevel models (Raudenbush & Bryk, 2002) have been used to study various aspects of interviewer effects. Early papers (Hox, 1994; Hox, de Leeuw, & Kreft, 1991) demonstrated the use of multilevel models in separating interviewer and respondent effects. Subsequently, multilevel models were used to assess the relative importance of geography and interviewer in contributing to measurement error (O’Muircheartaigh & Campanelli, 1998; Schnell & Kreuter, 2005), to study interviewer effects on non-response (Lipps & Pollien, 2011; Loosveldt & Beullens, 2014; Vassallo, Durrant, Smith, & Goldstein, 2015) and its components (Pickery & Loosveldt, 2002), to assess the role of non-response error variance and sampling error (rather than only measurement error) in generating interviewer effects (West, Kreuter, & Jaenichen, 2013; West & Olson, 2010), and to compare design-based and model-based methods to estimate interviewer effects in a complex sample (Elliott & West, 2015). One of the attractive features of multilevel models is that it allows the researcher to make predictions at the individual level as well as the group-mean level (Gelman & Hill, 2006). This naturally lends itself to a situation where we can try to estimate the effect of individual interviewers on measurement error (in our context, falsification). However, this seems an underutilized feature in the context of interviewer monitoring despite a few papers that have shown or called for its active use (Pickery & Loosveldt, 2004; West & Elliott, 2014). In this article, we demonstrate the use of multilevel models in making interviewer-specific predictions so as to spot interviewers who may be associated with falsifications. It is typically difficult to separate falsification from regular measurement error that arises from behaviors like interviewer probing. But the panel survey presented in this case study is unique in that interviewers are not supposed to play a role in data collection as substantive data are fully collected electronically without any manual intervention. In this situation, if the multilevel model—under some assumptions to be elaborated later—predicts significant interviewer effects, it would strongly suggest falsification rather than regular measurement error.
Survey and data
Television audience measurement (TAM) panels are a regular feature in many countries and produce “television ratings” (“TV ratings”) which are essentially estimates of TV viewing. These data are vital to the media and advertising industries; broadcasters use them to plan content, schedule programs, monitor program performance, and arrive at advertising rates, while advertisers use them to plan advertising so as to efficiently reach their target audiences. TV ratings, therefore, naturally become a common measure—a “currency”—by which advertisers and broadcasters are able to commercially engage with each other. In India, an organization called “TAM Media Research” (TAM India)—appointed initially by a Joint Industry Body (JIB) comprising Indian media and advertising organizations—operated the “currency” TAM panel between 1998 and 2015.
In its final year of operation, the panel had a sample size of 11,500 households (about 50,000 respondents) across 225 cities (primary sampling units [PSUs]) representing 98% of the estimated TV owning private households in urban India. Eight PSUs were self-representing (SR); non-SR PSUs were chosen with probability proportional to the estimated TV owning population within socio-cultural geographic regions nested within population strata. For various reasons, it was not possible to undertake probability-based household sampling within PSUs. Therefore, quota sampling was undertaken within strata formed using the following household variables: socio-economic class (SEC), type of access to TV (e.g., digital cable), size (number of members), preferred language of TV viewing, and geographic location within the PSU. One of the issues with quota sampling is that they can degenerate into a highly biased convenience sample; to ameliorate this, field supervisors were to give random start points to interviewers within a geographic area in a PSU. Following a serpentine order starting from the allotted random start point, interviewers used a fixed skipping pattern to contact households and administer a short demographic questionnaire (called a “listing”). The aim was to have the number of listings equal to about 15 times the design TAM panel size. These listings were accumulated, and based on the required quota, supervisors would sample households from these listings. Interviewers would then go back to the sampled households to try and recruit them onto the TAM panel. All individuals who were 4 years of age and above in a household who agreed to be part of the panel were recruited into the survey; tenure for a household was restricted to 3½ years. Households were incentivized using gifts such as household durables during major festivals. To ensure transparency, an independent body made up of global media experts called the TAM Transparency Panel (TTP) served as an ombudsman body between TAM India and the industry; a top-four consulting firm independently audited the fieldwork; and for the first 7 years of operations, the research design was vetted by the technical committee of the JIB.
An electronic device called the Peoplemeter was attached to each TV set of a recruited household. The Peoplemeter automatically collected data on which TV channels were tuned into by the household and at what time the tunings occurred. To capture information on individual viewing, a special remote control was provided to the household (one for each Peoplemeter) that had a unique button for each household member. A member pressed her button when she commenced viewing and pressed it again when she stopped viewing. All data were encrypted by the Peoplemeter and transmitted wirelessly to the central office for further processing. All in all, the data collection was designed to be free of human intervention. The interviewer’s role was to conduct listings, recruit households, train members in the proper use of the Peoplemeter equipment, update household information annually, disburse incentives, and generally serve as the one-point contact for the panel household. Interviewers were trained to avoid behavior that might influence respondents’ viewing, for example, they were to desist from conversations with respondents on specific TV channels or programs.
Three other distinguishing features of the TAM India data (apart from the automatic data collection) were the frequency of data reporting, data granularity, and data volume. Data were reported to the industry every Wednesday for the previous Sunday–Saturday data week. The implication was that there was only a narrow two workday window to process all data and run it through the required quality control processes. Data were reported for close to 600 active channels at a 1-min-level granularity across a range of demographics for the 50,000-individual sample base.
The scope for falsification
From the above data collection description, it might seem that there was no scope for any falsification as viewing data were generated and transmitted automatically without any human intervention. It might seem surprising then that detection of possible falsification was a core quality control task undertaken every week before data release. As stated earlier, the data were used as a “currency” by the media and advertising world. All else being equal, programs with higher viewership (more people watching and/or more time spent viewing on average) have a greater chance of garnering more adverting revenue. But getting higher viewership is not easy; the number of channels in India exploded from 67 channels in 1998 to about 600 channels in 2015 resulting in a very fragmented market where on average, 60% of the channels were estimated to be viewed by less than 1% of the population. Also, the population in India is extremely diverse—a mosaic of distinct languages, cultures, and religions—and some channels cater to a narrow segment of the population. Given this hyper-competitive scenario, there were increasing reports that unethical players (“intruders”) were adopting illegitimate methods to try to influence respondents’ viewing behaviors to boost viewership on some channels. There were two possible scenarios for this to occur depending on the nature of the intruder’s contact with the household:
Direct mechanism: Here, the intruder would directly contact the respondent. This would entail physically locating panel households and bribing them to view a certain channel or program more than others. As it would be a daunting task to locate panel households, help could be taken from parties such as local cable operators who would recognize Peoplemeter equipment when they would go to households to service their own equipment or collect monthly dues. Some interviewers also reported that they were being followed on their routes, ostensibly to locate panel households.
From an intruder perspective, the advantage in this mechanism is the direct contact with, and potential control over, the respondent. However, respondents are possibly less likely to acquiesce to such strangers’ requests. Moreover, falsification efforts would be nullified should the household inform the interviewer or the field supervisor; when being recruited, households were specifically told that they were to view TV completely in the way they wanted to and that they should report to TAM any contacts by external parties regarding their panel membership or any attempt to influence their viewing.
Indirect mechanism: Given the disadvantages in directly approaching a panel household, a more efficient choice for an intruder was to approach interviewers. A single interviewer would give access to many households—households that share a rapport with the interviewer and trust him. There were reports from some interviewers that parties were contacting them with bribes or threats to influence respondents. This also points to a special feature of the nature of falsification in this survey as compared with other general surveys: interviewers stood to be directly rewarded (or avoided getting hurt) by enabling the falsification of substantive data. Although many sincere interviewers reported these attempts back to management, it was impossible to assess the true extent of the problem though known falsification attempts were not widespread. Apart from the ethical nature of the problem, TV ratings data are critical to the industry and usage of these estimates depends on their credibility. Detection of possible falsification thus became a focal point in quality control.
Falsification detection methods
The usual methods to detect possible falsification were a combination of analysis of post-stratified aggregate viewing estimates and respondent-level checks. The former entailed looking at channels with, say, the average time spent being two standard deviations more than the last 12-week average, and then drilling down to possible causes. For example, an increase in viewing could be due to the telecast of a special program. TAM India has a division called AdEx India which monitors and records content for channels; analysts would check if the channels in question telecast special programs or changed their content in the time slot for which the increase in viewership was registered. Using this database, analysts were also able to check on-air promotion activity for programs. Respondent-level checks occurred via in-house software, a screenshot of which is shown in Figure 1.

Screenshot of the TAM India outlier detection software (some columns masked to protect confidentiality).
In Figure 1, we see time spent data over 21 days for six viewers of an English business news channel in a certain market. The software color-coded data points if they were more than preset upper control limits. The actual channel name in the software was coded (e.g., “i561” in the above screenshot) to remove the effect of analysts’ personal viewing preferences. Markets were also rotated across analysts every quarter.
We find that the second panelist has seen 120 min of the channel on Saturday (Day 7) in Week 8 of 2013 (the last data point in Row 2). Weekends typically see lesser business news viewing so this data point is atypical (though not extraordinary as viewers’ preferences are not always “rational”). The analyst would also notice that this panelist recently started watching more of this channel and also more frequently. Clicking on the data point would take the analyst into a detailed view of the panelist’s viewing across channels. The analyst would also check the respondent’s demographics; consistent heavy viewing of an English business news channel by a non-English speaking 10-year-old individual, for example, would be a cause of concern.
Analysts were operating under a narrow 2-day reporting window which made it impossible to conduct pre-data release checks such as some form of re-interviewing on households that showed suspicious viewing behavior. Therefore, respondents were classified by analysts into three classes: “Clear,” “Check,” and “Quarantine.” The “Check” class meant that the household’s data would be reported for the current week but its viewing behavior would be closely watched in the future. A “Quarantine” status was given to a household when its viewing was suspicious enough to not be reported for that week. A lot of discretion was exercised before classifying a household under this class so as to avoid superimposing personal ideas of what is “right viewing” and what is “wrong viewing” onto panelists’ viewing behavior. A “Quarantine” status for more than 2 weeks entailed a formal investigation by trained investigators who were members of the TAM India vigilance department headed by an ex-senior police officer. These investigations would involve interviewing respondents in a sensitive fashion; investigators avoided using channel names in the conversation unless absolutely required but tried to tactfully elicit responses from respondents on changes in their viewing behavior. If there was evidence that the household was compromised, it was removed from the panel and the interviewer also investigated. On the contrary, if there was no evidence of a falsification attempt, the household would be retained on the panel and feedback would be given to the quality control team to update internal records for future data checks.
Using formal models to detect possible falsification
Although the aggregate-level and respondent-level checks were useful, they were time and labor intensive. As the entire quality control activity had to be completed in 2 days, there was immense pressure on analysts; there were concerns of analyst fatigue leading to possible errors of judgment. There was also a need to make the process more objective, and explicitly take into account the possible falsification mechanisms. This motivated the use of a formal model-based method.
We demonstrate this method using the case of Channel XYZ in geographic reporting Market A. Channel names, the market name, and household and interviewer IDs have been masked to preserve confidentiality. Substantive data are used in an indexed form for the same reason. Market A had a total of 19 interviewers with workloads ranging from 7 to 34 households, with an average of 22.3 households. Most interviewers had workloads of more than 20 households. The viewing trend (based on poststratified data) for Channel XYZ and two of its competitors in Market A across 11 weeks is given in Figure 2; the trend is an index of average weekly time spent on the channels (called “viewing level”).

Viewership trend for Channel XYZ (solid line) and two of its competitors (dashed lines) in Market A across 11 weeks.
We see that as time progresses, Channel XYZ is on an upward trajectory; competitor channels’ viewerships have also grown, but not as much as Channel XYZ. The gradual viewing increase for Channel XYZ poses a problem for the conventional data checking process described in “Falsification detection methods” section above; the data “creep up” on an analyst, and at the end of 11 weeks, Channel XYZ has registered a 55% increase in viewing. We see a viewing spike for Channel XYZ in Week 8; a check on the AdEx database found that there was no change in the content on Channel XYZ during that week.
Multilevel-model specification
The nested structure of the data naturally lent itself to a multilevel longitudinal model; longitudinal viewing level data across weeks (
We use a hierarchical model specification (Raudenbush & Bryk, 1986) that clearly reflects the three levels in the data:
Level 1: Longitudinal data for Channel XYZ for 11 weeks (t = 1, 2, . . . ,11)
Here,
Level 2: Household, i
We see that the household-specific intercepts
In the above equations,
As part of
Past research indicates that model-based interviewer variance estimates substantially decrease when geographic variables, such as the proportion of population below the poverty line, are included (West et al., 2013). However, such geographic contextual data were unavailable in our case.
Level 3: Interviewer, j
Here, we have
In specifying interviewers as a random effect, we are conceptualizing the interviewers in the survey to be drawn from a larger population of interviewers. We did not add any interviewer-level covariates (such as interviewer experience) at this level as our interest is not in explaining the interviewer-specific intercepts but in identifying possibly problematic interviewers.
Unlike the typical use of regression models, we are not interested in explaining or predicting a substantive response. Our interest is focused on realizations of the household-level random effects (
The lme4 package (Bates, Mächler, Bolker, & Walker, 2015) in R (R Core Team, 2014) was used to fit the model; restricted maximum likelihood (REML) was used using a Laplace approximation for the likelihood evaluation. Household random intercepts were added first, followed by household random slopes, interviewer random intercepts, and, finally, interviewer random slopes. At each step, a non-zero variance of the random effect was tested using a 50:50 χ2 distribution (Self & Liang, 1987; Stram & Lee, 1994) at a 5% significance level.
Identifying potentially problematic panel interviewers or households was accomplished by extracting the Empirical Best Linear Unbiased Predictors (EBLUPs) for those random effects whose variance term was statistically significant. Two types of prediction intervals were computed for each EBLUP to test whether it was statistically greater than zero (i.e., corresponding to one-sided tests): nominal 95% prediction intervals and prediction intervals computed after a Bonferroni correction was applied to the nominal 5% type 1 error rate.
Results
We present results for the interviewers first as interviewer-led falsifications were more likely than direct household-led falsifications (as discussed in the “The scope for falsification” section). We failed to reject the null hypothesis of interviewer variance in slopes (χ2 = 0.65; p = .91). Variance of the random interviewer intercept was found to be significant (χ2 = 17.4; p = .0003); EBLUPs are presented below in Figure 3.

Sorted EBLUPs of interviewer random intercepts for Channel XYZ in Market A.
EBLUPs of Interviewers 319 and 323 are found to be statistically greater than zero based on both standard and Bonferroni-adjusted prediction intervals. To see how this finding relates to the actual viewing data, Channel XYZ’s viewing trend for households across interviewers was plotted in Figure 4.

Viewing level (vertical axes) trends of Channel XYZ for households belonging to different interviewers.
EBLUPs of the household random effects are presented in Figure 5. A visual inspection of Figure 5 led to analyzing outlier households (darker points at the upper end of the tails in Figure 5) separately in Figure 6. Eight households were isolated on the basis of EBLUPs of random intercepts, and seven households on the basis of EBLUPs of random slopes.

(a) Sorted EBLUPs of household random intercepts and (b) household random slopes for Channel XYZ in Market A.

(a) Sorted EBLUPs of household random intercepts and (b) viewing trend of these households for Channel XYZ in Market A. Corresponding plots for random slopes are in (c) and (d). Numbers on the vertical axes in (a) and (c) are household IDs with the interviewer IDs of those households alongside in parenthesis. Vertical scales in (b) and (d) have the same range for ease of comparison.
Figure 6(a) shows that all eight random intercept EBLUPs for the outlier households are significant. Figure 6(c) shows that three of the seven random slope EBLUPs have their Bonferroni-corrected confidence intervals just including zero. However, given that the Bonferroni correction is conservative (Faraway, 2005), these households were also included for further analysis and investigation.
Although there are no common household IDs between Figure 6(a) and (c) (in which case they would have been the most suspected households), Figure 6(b) shows that some households (chosen on the basis of EBLUPs of random intercepts) also show fairly large increases in viewing. Viewing changes are registering on a higher base for these households resulting in their getting picked on the basis of their random intercept EBLUPs rather than their random slope EBLUPs. But the increase in viewing for these households is very apparent in Figure 6(b) and these households can be prioritized for investigation. Interestingly, these households register a change in viewing around Week 7, which is when we see outlier households based on EBLUPs of random slopes registering their increases in viewing as well (Figure 6[d]). As noted earlier in the “Using formal models to detect possible falsification” section, there was no change in content for this channel, so this coordinated increase in viewing (across outlier households) is suspicious. Finally, Figure 6(a) and (c) shows that about half the outlier households belong to the two outlier interviewers.
These results are not a confirmation of falsification by the identified interviewers. For example, while unlikely, the areas under these two interviewers may exactly be those where Channel XYZ had undertaken heavy marketing promotions. It is also likely that intruders may have succeeded in locating panel households in only the areas corresponding to the two interviewers. Therefore, formal investigations (as described in “Falsification detection methods” section) were launched by TAM India to confirm whether the households were falsified and the identified interviewers involved. The investigations covered all the households from Figure 6 and, apart from these households, as many households as possible from Interviewers 319 and 323. The investigations supported model-based predictions. Although legally admissible evidence could not be garnered, there were strong indications of respondents’ viewing being influenced. Accordingly, the suspicious homes were dropped from the panel and the two interviewers put on special scrutiny. No further issues originated from this market.
Discussion
Assuming that the model satisfactorily approximates an interpenetrated design, and assuming sufficient power, we ideally should not have any interviewer effect, that is, we should fail to reject the hypothesis that
This is because in our survey, interviewers do not have any role in the data collection and therefore should not have any effect on respondents’ viewing data. However, the fact that we found a significant interviewer random effect implied that there might be specific interviewers who need a closer look. The plot of EBLUPs of interviewer random effects in Figure 3 pointed to two possibly problematic interviewers. Household-specific plots showed that many heavy viewing households are concentrated within these two interviewers. On-ground investigations supported our model-based inferences.
The EBLUPs of random effects are shrinkage estimators (Hox, 2010; McCulloch, Searle, & Neuhaus, 2008; Robinson, 1991), and as implied in the “Using formal models to detect possible falsification” section, interviewer workloads were imbalanced. EBLUPs take into account the fact that interviewer-specific intercepts will be predicted imprecisely for interviewers with small workloads, to compensate for which they are “shrunk” toward the overall intercept. This would have the effect of some masking of actual differences between interviewers. Therefore, it would also be important to look at the household-level EBLUPs in addition to the interviewer-level EBLUPs to detect potential problems. Another reason for looking at household-level EBLUPs is to handle situations where, for illustration’s sake, many interviewers have been tapped to influence just one household each in their workload. A third reason to look at household-level EBLUPs is to sharpen investigation efforts; having known that two interviewers may be problematic, it would also be useful to know which households in their workloads could have been influenced. The household-level controls in the model help to isolate interviewer effects by approximating an interpenetrated design. They also enable us to spot potentially problematic households by ironing out demographic profile differences among households.
If an analyst was shown only Figure 4 as is the case in common data checking methods, she might have picked up Interviewer 318 too as problematic. But Figure 3 shows that this interviewer is the one with the least positive (and statistically insignificant) EBLUP; the model controls for several household characteristics simultaneously while pointing to a particular interviewer, which the analyst cannot do when she looks at Figure 4 alone. This demonstrates the effectiveness of the model-based method over the mere visual inspection of the data.
We did not use the poststratification weights in this analysis. Although methods to incorporate survey weights in multilevel models exist (Pfeffermann, Skinner, Holmes, Goldstein, & Rasbash, 1998), making inference about interviewer effects on viewing levels averaged across the household population distribution is not the goal of this analysis. Rather, we are interested in determining impacts across the population represented by those actually sampled as the key measure of falsification, suggesting the use of unweighted model estimates.
Although singling out households from Figure 5 was a subjective decision too (based as it was on visual inspection), it was based on a formal model accompanied by the computing of confidence intervals making the process much more objective. Also, every analyst uses an implicit personal model when marking a household as one that needs to be checked or quarantined. The multilevel method used here makes the model explicit and uniform across analysts while scientifically controlling for various variables and is, therefore, a more objective method. The ability to quickly parse through hundreds of channels makes this method very efficient, saving a large proportion of analyst time. The demographic controls approximating an interpenetrated design can remain constant across runs, and all that is needed is to test the variance components and plot the EBLUPs of significant random effects. Running the model for the present case along with producing the relevant figures took about 3 s on a Windows laptop with a 2.9 GHz Intel i7 processor and 8 GB RAM. Data can be updated every week and models run on them. There could be a concern that outlier households/interviewers may not show up for a current week unless very extreme as the models take into account all past data. We therefore also recommend running these models for the current week without the time component.
When on-ground investigations show that households are compromised, we recommend retaining them in the panel but taking care not to report their data. All else being equal, by not reporting the compromised households’ viewing, estimates for the offending channel would have lesser magnitude as compared with previous weeks. Smart intruders typically track data for the effectiveness of their falsification efforts and would, therefore, observe this drop in viewing estimates. Consequently, they may give instructions to respondents (or ask interviewers to do so) to increase levels of viewing. This will be captured in the data (but not part of the reported data), thus giving valuable insight into falsification mechanisms. In such circumstances, it is important that the survey organization document these actions and make them transparent to parties such as certified third-party auditors.
As pointed out by a reviewer, while falsification occurs at a household level, it still requires a household member (or a group of members) collaborating with a falsifying interviewer to decide which members to log-in; sophisticated operators would know that logging-in all members of the home would raise suspicion. Therefore, apart from extending the multilevel analysis to the individual level, survey managers would also benefit from using a version of Cover Analysis (Danaher & Sharot, 1994) to identify log-ins that may signal falsification. “Cover” is defined as the time spent by a member in a panel household divided by the sum of minutes spent by all members of the household. A model can be fit using “Cover” as the outcome variable with household and individual variables as inputs. Assuming a well-fitting model, panelists with observed values more than a certain threshold away from the predicted values (say 1.5 times the predicted value) can be singled-out for investigation.
The methods in this article can be implemented for cross-sectional surveys as well to act as early-warning mechanisms. One way this can work is to wait for a certain proportion of fieldwork to be completed. Multilevel models are then fit to each item (or at least important items) of the questionnaire using data collected until then. This would enable the analyst to identify specific interviewers contributing to measurement error (such as seen in Figure 3) for each item. Patterns from this analysis can be used to inform interviewer retraining, for example, the survey manager may find that an interviewer is consistently identified for sensitive items, which means that this interviewer may need more training to administer such questions.
We recognize some of the drawbacks in the method and its application. First, though we included households belonging to other interviewers, our focus was on Interviewers 319 and 323 as these appeared to be the most problematic. However, even though associated with “insignificant effects” in Figure 3, investigating Interviewers 327 and 324 more extensively would have provided a measure of model validation (we had only two households from these two interviewers that were actually investigated). However, this was not undertaken due to resource constraints. In general, survey managers should consider looking at least a little beyond purely “significant” results. Second, we assume that the model truly approximated interpenetration by controlling for various demographic variables. But a phenomenon like a TV channel’s viewing could be associated more with psychographics than plain demographics; we have no way to verify whether controlling for demographics also led to successfully controlling for other aspects that influence TV program viewing preferences. Third, the models are suited for situations with a “large” number of interviewers. As with any discussion on sample size, there is no fixed rule to what constitutes a “large” number of interviewers but we may run into estimation problems when the number of interviewers is small. Fourth, one requires some basic heterogeneity in interviewer assignments to make the required interviewer-led inferences. As an extreme example, if each interviewer had a unique respondent profile assigned to him, then interviewer effects would be confounded with respondent demographics. A similar confounding can occur between geographic and interviewer effects. Fifth, it may not be possible to conduct model diagnostics in a production environment with tight turnaround times. This may not be as critical an issue in our case as we are not making inferences for a substantive variable. Evidence is mixed on whether the assumed distribution of the random effects affects the EBLUPs of the random effects (McCulloch & Neuhaus, 2011a; McCulloch & Neuhaus, 2011b; Skrondal & Rabe-Hesketh, 2009), but in any case, judging the true distribution of the random effects based on the EBLUPs can be misleading (McCulloch & Neuhaus, 2011a). One option is to judge the sensitivity of predictions under different distributional assumptions using software like the hglm package (Ronnegard, Shen, & Alam, 2010) in R (R Core Team, 2014). Finally, it might be that use of a two-component mixture distribution for the random effects components might substantially increase power for falsification detection by allowing for a (presumably smaller) component of random effects to be centered at a positive mean corresponding to the “average” effect of falsifiers. Given that a small number of falsifiers are likely to exist in a sample, the practical value of such an approach should be first assessed by an extensive simulation study. In addition, the limited software to fit three-level models with mixture components will pose problems as well, although rapid development in this area may make this a short-term difficulty.
Conclusion
Most respectable surveys conduct some sort of quality control before data are released. However, many of the traditional procedures (such as presented in “Falsification detection methods” section) are time-consuming and effort-intensive. With the increasing demand for quicker data, surveys will find themselves under pressure to release data faster, requiring efficient and effective quality control procedures. Model-based methods such as those described by this article can be valuable in this regard.
Case studies are rare in the survey falsification literature for at least three reasons. First, some organizations are hesitant in disclosing their internal detection processes, as they view them as a competitive advantage. Second, organizations are concerned that publishing such case studies might entail disclosure risk. Third, publishing case studies is seen as an acknowledgment that falsification does occur within the system; there may be concern that this may be seen as a weakness by sponsors and lead to uncomfortable questions on the prevalence of falsification and its impact on data. The first challenge can be overcome by encouraging organizations to sketch the general scientific method without revealing details that are proprietary. Similarly, disclosure risk concerns can be overcome by masking names of sample towns and items, and using indexed substantive data. Finally, falsification is the proverbial “elephant in the room.” In our experience, it is better to openly acknowledge that a problem may exist and attempt to tackle it sincerely. Apart from helping practitioners design their quality control systems, our hope is that sharing this case study also encourages other market research and survey organizations to report their experiences with falsification. As more organizations do this, sponsors will appreciate the complexity of the problem and there will be a greater possibility of a concerted effort at jointly tackling what is a serious ethical and research issue.
Footnotes
Acknowledgements
The authors are very grateful to Mr. L.V.Krishnan, CEO of TAM India, for permission to publish this research. Thanks to Sajid Qureshi and Kapil Vakharia at TAM India for help with data management.
Declaration of conflicting interests
The author(s) declared following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: Sharan Sharma was an employee of TAM India when this research was conducted.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
