
Other
Select search scope: search across all journals or within the current journal


In recent years, the random noise method has been gaining wider use in National Statistical Institutes (NSIs) to protect respondent data from disclosure. The random noise method takes a micro-data approach to disclosure limitation: a multiplier (noise factor) is applied to each unit prior to tabulation – thus, guaranteeing that different tabulations, from the lowest to the highest level, are consistent. In this paper we evaluate two different random noise models in an applied context. Our analysis suggests that both the Basic Noise Model (BNM) and the Alternate Noise Model (ANM) are unsatisfactory for protecting smaller units. To overcome this difficulty, we developed a(third) Mixed Perturbation Model (MPM) that combines the use of multiplicative noise to protect large units with the use of synthetic models to protect the smaller units. To accomplish this, we constructed a hybrid model (first logistic and then linear) to generate the synthetic data. Results indicate that the mixed approach performs better than either of the other two models, both in terms of reliability and disclosure limitation, although it too has weaknesses. Hence, areas for future study remain that our research suggests might be tackled next.



If data availability were a simple problem, it would already have been resolved. In this paper, I argue that by viewing data availability as a public good, it is possible to both understand the complexities with which it is fraught and identify a path to a solution.
Data openness is an important issue for statistical systems in many countries, particularly in the developing world. Data do not become¡°open¡± overnight, even if governments so desire. Openness has several components. First, one has to deal with legal issues of openness. Then there are organizational and technical issues, having to do with compiling and presenting data in open formats. In addition, underlying data quality issues surface when data become more open. Furthermore, there are often conflicting interests advocating for or trying to limit data openness, within the government, in civil society, and in the private sector. Therefore, opening databases cannot be accomplished by a simple act of "good will" on the part of government; it entails a lot of preparation and the balancing of many interests. While large, multilateral organizations have recently become notable advocates of open data and, more broadly, "open government," their interests and practical capacities are often limited by their mandate: some may be interested only in economic statistics, or national-level health statistics, for example. And large international agencies are often unable or unwilling to engage with civil society organizations or other interest groups, who are potential users and producers of data. Still there is a need for their financial support for complex reforms. But these agencies may themselves be limited in their authority or capacity. Furthermore large multilateral organizations may not be able to engage civil society in some countries due to political antagonisms or other circumstances. Therefore there is a niche for non-governmental organizations, bringing international experience adaptable to local conditions to serve as technical resources and trusted advisers to governments and to act as go-betweens with foundations and international agencies that are prepared to support open data reforms.

National Statistical offices (NSOs) create official statistics from data collected from survey respondents, government administrative records and other sources. The raw source data is usually considered to be confidential. In the case of the U.S. Census Bureau, confidentiality of survey and administrative records microdata is mandated by statute, and this mandate to protect confidentiality is often at odds with the needs of users to extract as much information from the data as possible. Traditional disclosure protection techniques result in official data products that do not fully utilize the information content of the underlying microdata. Typically, these products take the form of simple aggregate tabulations. In a few cases anonymized public-use micro samples are made available, but these face a growing risk of re-identification by the increasing amounts of information about individuals and firms available in the public domain. One approach for overcoming these risks is to release products based on synthetic data where values are simulated from statistical models designed to mimic the (joint) distributions of the underlying microdata. We discuss recent Census Bureau work to develop and deploy such products. We discuss the benefits and challenges involved with extending the scope of synthetic data products in official statistics.
Distributions of business data are typically much more skewed than those for household or individual data and public knowledge of the underlying units is greater. As a results, national statistical offices (NSOs) rarely release establishment or firm-level business microdata due to the risk to respondent confidentiality. One potential approach for overcoming these risks is to release synthetic data where the establishment data are simulated from statistical models designed to mimic the distributions of the real underlying microdata. The US Census Bureau's Center for Economic Studies in collaboration with Duke University, the National Institute of Statistical Sciences, and Cornell University made available a synthetic public use file for the Longitudinal Business Database (LBD) comprising more than 20 million records for all business establishment with paid employees dating back to 1976. The resulting product, dubbed the SynLBD, was released in 2010 and is the first-ever comprehensive business microdata set publicly released in the United States including data on establishments' employment and payroll, birth and death years, and industrial classification. This paper documents the scope of projects that have requested and used the SynLBD.
In most countries, national statistical agencies do not release establishment-level business microdata, because doing so represents too large a risk to establishments' confidentiality. Agencies potentially can manage these risks by releasing synthetic microdata, i.e., individual establishment records simulated from statistical models designed to mimic the joint distribution of the underlying observed data. Previously, we used this approach to generate a public-use version – now available for public use – of the U.S. Census Bureau's Longitudinal Business Database (LBD), a longitudinal census of establishments dating back to 1976. While the synthetic LBD has proven to be a useful product, we now seek to improve and expand it by using new synthesis models and adding features. This article describes our efforts to create the second generation of the SynLBD, including synthesis procedures that we believe could be replicated in other contexts.
One major criticism against the use of synthetic data has been that the efforts necessary to generate useful synthetic data are so intense that many statistical agencies cannot afford them. We argue many lessons in this evolving field have been learned in the early years of synthetic data generation, and can be used in the development of new synthetic data products, considerably reducing the required investments. The final goal of the project described in this paper will be to evaluate whether synthetic data algorithms developed in the U.S. to generate a synthetic version of the Longitudinal Business Database (LBD) can easily be transferred to generate a similar data product for other countries. We construct a German data product with information comparable to the LBD – the German Longitudinal Business Database(GLBD) – that is generated from different administrative sources at the Institute for Employment Research, Germany. In a future step, the algorithms developed for the synthesis of the LBD will be applied to the GLBD. Extensive evaluations will illustrate whether the algorithms provide useful synthetic data without further adjustment. The ultimate goal of the project is to provide access to multiple synthetic datasets similar to the SynLBD at Cornell to enable comparative studies between countries. The Synthetic GLBD is a first step towards that goal.