This is a resume of salient aspects of Warner's pioneering work on randomized response techniques and related work over the fifty years since its inception, subject to the author's interest.
Select search scope: search across all journals or within the current journal
This is a resume of salient aspects of Warner's pioneering work on randomized response techniques and related work over the fifty years since its inception, subject to the author's interest.
In a social survey an unequal probability sample is supposedly at hand. The problem of utilizing it to estimate the finite population mean of a stigmatizing variable has a known solution admitting assessment of accuracy level. Here we examine improvement by Empirical Bayes procedure, exploiting available data on a correlated variable while using a randomized response technique to gather data on the main variable of interest.
Traditional survey techniques on human populations are expected to give poor results when the issue under investigation is sensitive or stigmatizing. Various sources of nonsampling error, in particular nonresponse and misleading answers, are serious threats to the validity of the conclusions. Indirect questioning techniques offer a remedy to this problem. Randomized response, introduced by Warner [14], has the lion's share in indirect questioning, but an alternative, the item count technique, is popular among social scientists. Although the original version introduced by Raghavarao and Federer [10], Miller [8], and Miller et al. [9] is easily understood by respondents and can be incorporated in structured questionnaires, it does not fully protect the privacy of the participants. In this paper, where our main priority will be the protection of privacy, we present a modified version of the item count technique.
In estimating the proportion of people bearing sensitive matters like habits of tax evasion, drunken driving, etc. in a given community, it is difficult to obtain trustworthy data through direct queries. To overcome this difficulty, Warner [29] introduced randomized response techniques to estimate the proportion of people bearing such a stigmatizing or sensitive characteristic in a given community. Since then, several researchers have extended and applied this technique in various ways, for instance, Greenberg et al. [13], Horvitz et al. [16], Mukerjee [23], Ljungqvist [20], Christofides [10], Singh and Grewal [26], and many others. In many areas of the RR-related activities, the sample selection is traditionally by simple random sampling (SRS) with replacement (WR). Keeping in mind that the general large scale sample surveys usually involve the unequal probability sample selection, subsequently many researchers have enriched the RR-related literature by extending in unequal probability sampling (see Chaudhuri et al. [5], Dihidar [12], Chaudhuri and Dihidar [6]). Hanurav [14], Rao [25], Chaudhuri and Arnab [4] had compared some unequal probability sampling strategies for estimating the population mean of a quantitative variable in direct surveys under a super-population model. In this paper, we consider the problem of estimating sensitive population proportion by unequal probability sampling using the pioneering randomized response techniques due to Warner [29], Mangat and Singh [22], Christofides [9], Chaudhuri and Mukerjee's [7] forced response model, Kuk's model [18], Singh and Joarder's [27] unknown repeated trial model and compare the Horvitz-Thompson's [17] and Murthy's [24] strategies under the super population model as proposed in Lanke [19]. It is shown that under this model-cum-design based approach, Murthy's strategy performs better than Horvitz-Thompson's strategy.
Randomized response technique is an effective research method that is used to estimate the proportion of a population that possesses a sensitive characteristic, such as tax evasion, abortion, or drug abuse. In this paper, two recently developed randomized response models using two decks of cards are extended to the case of stratified random sampling. Under proportional allocation or optimal allocation for fixed sample sizes, it is shown that the proposed stratified estimators are always more efficient than their counterparts in simple random sampling. Comparing the two proposed models, it is found that one of the models can be adjusted to be more efficient than the other, and that this model can increase the respondents' cooperation. Hence this model is referred to as the improved model. For unknown strata sizes, the double sampling method for stratification is applied.
In this paper, we suggest the Bayes linear estimator (BLE) for randomized response model (RRM) to improve the efficiency of RR estimators, only using the first and second prior moments. The randomized response model is an indirect questioning technique used to protect the privacy of respondents in a survey regarding a sensitive characteristic. Meanwhile Bayes linear estimation is useful for parameter estimation compared to the typical Bayesian method because it only uses the first and second prior knowledge of the variable of interest. Also, it has an advantage of robustness with the distribution.
We suggest the Bayes linear estimators for the two-stage and the stratified RRM and find the optimal sample size to minimize the Bayes risk for the stratified RRM. Also, we show the difference in efficiency between the Bayes linear estimators and the typical non-Bayesian RR estimators by simulation study.
Randomized response (RR) was introduced as a technique for protecting respondents' privacy in survey interviews regarding sensitive characteristics. In recent years, the basic RR ideas have been used and extended in other contexts. We discuss usage and recent advances of RR in confidentiality protection and in privacy preserving data mining. We discuss important differences between RR surveys and RR for confidentiality protection. In particular, for confidentiality protection, the data may be used to choose suitable randomization probabilities, but doing so renders well known inferences derived for RR surveys inapplicable. We examine one privacy breach criterion in data mining and propose a new privacy guarantee and a method for its achievement. We also discuss several new challenges and open problems for future research.
The crux of this paper is to estimate the mean of the number of persons possessing a rare sensitive attribute based on the Singh et al. [1] randomization device by utilizing the Poisson distribution in stratified sampling. This study also deals with the extension of the estimation reported by Singh and Tarray [2] using a Poisson distribution and an unrelated question randomized response model reported in Singh et al. [1]. In stratified sampling, the estimators are proposed when the parameter of the rare unrelated attribute is known and also when it is unknown. It is shown that the proposed models are more efficient than the model given by Lee et al. [3] in both cases, that is, when the proportion of persons possessing a rare unrelated attribute is known and that when it is unknown. When the sizes of the stratified populations are not given, other estimators are suggested using stratified double sampling. Properties of the proposed randomized response model are studied and recommendations are made.
In this article, some new randomized response models have been proposed. Properties of the proposed randomized response models have been studied. The proposed models are found to be more efficient than the randomized response models studied by Himmelfarb and Edgell [1], Gjestvang and Singh [2] and Singh [3] under certain realistic conditions. Numerical illustrations are also given in support of the present study.
We consider the problem of unbiased estimation of a finite population proportion and compare the unequal probability sampling strategies due to Hansen-Hurwitz [9], Horvitz-Thompson [12], Rao-Hartley-Cochran [18] and Midzuno-Sen [16,20] under a super-population model. It is shown that the model expected variance is least for the Midzuno-Sen [16,20] strategy both when these sampling strategies are based on data obtained from (i) a direct survey and (ii) a randomized response (RR) survey employing some RR technique following a general RR model.
In sample surveys it is often difficult to obtain true responses to questions of a personal or sensitive nature, such as questions regarding savings, drug use, or extramarital affairs. To avoid providing the requisite information, or to avoid embarrassment, some respondents may refuse to give answers or may give false answers. Thus the estimates obtained from a direct survey on such topics would be biased and inferences drawn from these would be erroneous. In order to solve this problem a number of randomized response techniques (RRT), pioneered by Warner [14], have been developed. Summaries of such techniques have been made by Chaudhuri and Mukherjee [4], Mukhopadhyay [11], and Chaudhuri [3]. Here we shall find an optimal estimator of the population total of a sensitive character y in surveys using randomized response techniques by an application of optimal estimating functions.
Warner [10] pioneered the randomized response (RR) technique for collecting information on a sensitive qualitative characteristic. Following Rao [7] and Chaudhuri and Stenger [6] an attempt has been made to derive an exact expression for the mean square error of ratio estimator, regression estimator, separate and combined ratio estimator, separate and combined regression estimator in stratified random sampling and the generalized regression estimator in case of randomized response surveys dealing with a sensitive quantitative character, for example, expenditure on drinking alcohol, additional income earned over usual salary to hide income tax in the income tax return, expenditure on gambling, number of induced abortions, income earned through prostitution, expenditure incurred towards payment of prostitutes in a red light area and some such others which we do not want to disclose even to our wives. Also an exact expression for an unbiased estimator of the mean square error of each of the above estimators is worked out.
Item Count Technique was added as a procedure of estimating a finite population proportion of any sensitive characteristic so as to increase the level of protection of privacy of the respondents as compared to the Randomized Response Techniques (RRT's) by Warner [14] and the follow-ups. Raghavarao and Federer [13], Miller [11], Miller, Cisin and Harrel [12] introduced the Item Count Technique also known as List Experiments or the Block Total Response or the Unmatched Count Technique. Chaudhuri and Christofides [5] have extended the Item Count Technique for estimating finite population proportion of any sensitive attribute as well as finite mean or total of any sensitive quantitative variable from a sample drawn using a general sampling design whereas the original method was needlessly restricted to simple random sampling. The Chaudhuri and Christofides [5] version of Item Count Technique for estimating a finite population mean or total of any sensitive quantitative characteristic has an advantage over those by Eichhorn and Hayre [7], Gupta, Gupta and Singh [8] and Huang [10] in the sense that the population means of the innocuous quantitative characteristics used need not be known. A drawback of the method by Chaudhuri and Christofides [5] is that it requires the selection of two independent samples costing more time and money. Hence a revised method is presented in this paper.
A survey was conducted on the students of a certain University to determine the prevalence of various risky behaviors. The data was collected using direct method (DR) and randomized response (RR) method. It is found that RR method yielded higher estimates of the prevalence. The estimated standard errors based on RR surveys did not always produce expected results that the estimates should be higher than that of the DR surveys.
Here we consider an interesting example of a randomized response model and see how diverse statistical techniques can be used in estimating parametric functions of a sensitive variable. We compare the suggested estimator under the proposed model to some well known estimators, including Warner's celebrated estimator.
Untruthful answers and nonresponse might cause serious problems in sample and population surveys on sensitive variables. In honor of the 50 year jubilee of Warner's basic idea in the field of randomized response techniques, this particular strategy is extended to general probability sampling with arbitrary sample inclusion probabilities. The statistical properties of the derived estimator for the relative size of an interesting subpopulation are presented in the context of data imputation. Moreover, the theory is extended to the realistic case that some of the respondents might be willing to answer on the direct question. The effect of a certain behavior of the interviewees on the statistical properties of the estimator is discussed.

