
Editorial
Select search scope: search across all journals or within the current journal


The growing importance of statistical evidence, data and information for political decisions is reflected in the handy and popular formulation ’Data for Policy’ (D4P). Under this cover, well-known guiding themes, such as the modernisation of the public sector, or evidence-informed policy-making, are led to new solutions with new technologies and infinitely rich data sources. Data for Policy means more to official statistics than just new data, techniques and methods. It is not least a matter of securing an important function and position for official statistics in the Policy for Data of the future. In order to justify this position, it is necessary to have a clear understanding of the tasks of official statistics for the functioning of (democratic) societies, with a view to how these tasks have to be reinterpreted under changing conditions (above all because of digitisation and globalisation).
Declining public trust in official statistics and indicators is frequently highlighted as a key obstacle to reasoned debate on policy options and governance choices. The potentially harmful impacts of Big Data and an alleged “post-truth” era have further accentuated such concerns. To remain trusted and credible, statistical institutions must safeguard their authority as sources of independent and scientifically sound indicators, while at the same time being innovative, to ensure the relevance of the indicators. However, this article argues that, in addition to this trust-building work, embracing mistrust and distrust is essential if indicators are to be relevant and influential. By unpacking the notion of trust, the article illustrates ways in which mistrust and distrust can serve as resources rather than mere threats to the credibility and authority of official statistics. For further empirical work, a conceptual framework consisting of three dimensions of trust and a distinction between mistrust and distrust is proposed and illustrated with concrete examples from indicator work. The conclusions suggest ways for statistical institutions to adjust their strategies so as to maintain trust via a more nuanced understanding of the multiple dimensions of trust, mistrust and distrust.
As societies evolve through the data revolution, it is important that National Statistical Offices (NSOs) continue to devote efforts to fine-tune their approaches to maintain the vital trust they need to operate successfully. To do this, they must constantly re-invent themselves to remain relevant to new data needs and keep up with their high-quality standards. With new sources of information surfacing both in the public and private spheres, options are multiplying for NSOs to design new ways to gather and grow the data into information. As this is happening, new practices and new issues are emerging throughout the data life-cycle process. Operating beyond the sample survey paradigm, NSOs see themselves confronted with the need to anchor their new approaches in solidly defined and defendable frameworks. Further, as new data themes, methods and sources are considered, transparency becomes a central issue. Using Statistics Canada’s data life-cycle management model, this paper illustrates how the scientific approach can be leveraged to make transparency more explicit both in projects and management.
Statistical disclosure control aims neither at revealing the identity of an individual, nor at revealing characteristics of individuals, households or companies that are confidential or personal. Primary statistical secrecy concerns information one can directly assess, whereas secondary statistical secrecy concerns information that a user could deduce indirectly by recombining and crosschecking all the disseminated data.
In the case of spatial data disseminated according to several geographical partitions, it is possible to combine and intersect the geographical areas in order to derive information on new and smaller areas. The differencing technique, which consists in subtracting the value of two overlapping areas, can lead to a breach of confidentiality.
We have developed a method for dealing with geographical differencing problems by detecting individuals located in small overlapping areas and whose personal information can therefore be disclosed. Modelling the data into a graph structure enables focusing on relevant geographical regions. The originality of the method resides in reducing the graph size and complexity. We applied the method to French income tax data composed of 27 million households dispatched on the 150 000 square cells and on the 35 000 administrative units. The results show that 10 000 households are at risk for disclosure.
Interview falsification has been occurring for decades on different levels but is still often insufficiently detected and treated in practice. This is a severe issue as interview fabrication can have major effects on final estimates.
This article aims to draw attention to this important topic, with focus on falsification by telephone (CATI) interviewers, and to provide National Statistical Offices with a base for a data-driven, mostly automatic tool written in the statistical software R to detect interviewers with irregularities in their answers or their time patterns. The further application of traditional methods, e.g. recontact or test calls, is far more effective with such focused samples than with random samples.
As revealing conspicuous interviewers sometimes feels like detective work, this article intentionally has a little detective side story. Together we will look at the means of evidence, investigate it with various methods, filter out the conclusive evidence and conclude how to close the case.
Stats NZ’s Integrated Data Infrastructure (IDI) is a linked longitudinal database combining administrative and survey data. Initially, the IDI contained a small number of administrative datasets from key government agencies, which generally contained good quality identifying information such as names and date of birth. As a result, the methodologies developed to link these datasets together relied heavily on these variables, and yielded high link rates while maintaining good quality links. When survey datasets were later added, the link rates achieved were lower than that of administrative datasets, due to poor quality names and the underutilisation of geographic information. This indicated there were improvements to be made to the linking methodology used to link survey data in the IDI. Stats NZ underwent extensive consultation with the research community on their requirements for the expansion of the IDI (IDI2). A key finding from the consultation was that researchers wanted improved survey linkage. This paper outlines how the address history of individuals were used to increase the link rate of surveys in the IDI.

In this contribution we outline the concept of Trusted Smart Statistics as the natural evolution of official statistics in the new
Recent years have seen dramatic changes in sources of data, amounts of data, availability of data, frequency of data, and types of data. Along with advances in data analytic technology these changes have opened up huge possibilities for improving the information content and timeliness of official statistics. in this paper we characterise such “smart statistics”, examining their potential benefits and the obstacles that must be overcome if they are to be trusted and relied upon. In particular, we list eight specific recommendations which we believe producers of smart statistics should adhere to if the full potential for economic and social benefit is to be achieved.
National statistical institutes are using frameworks to organise and set up their official statistical production, e.g. GSBPM. As a sequential approach of statistical production, GSBPM has become a well-established standard using deductive reasoning as analytics’ paradigm. For example, the first GSBPM steps are entirely focused on deductive reasoning based on primary data collection and are not suited for inductive reasoning applied to (already existing) secondary data (e.g. big data resulting, for example, from smart ecosystems). Taken into account the apparent potential of big data in the official statistical production, the GSBPM process needs to adapted to incorporate both complementary approaches of analytics (i.e. inductive and deductive reasoning) and, for example, through the usage of, for example, data-informed continuous evaluation at any GSBPM step. This paper discusses the limitations of GSBPM with respect to the usage of big data (using inductive reasoning as analytics’ paradigm), and also with respect to trusted smart statistics. The authors give insights on how to augment and empower current statistical production processes by analytics, and also by (trusted) smart statistics. In addition, the paper also highlights challenges and opportunities that should be addressed to embrace this major paradigm shift.
Privacy by design (PbD) as an approach to systems engineering has been conceptualized for more than a decade. It has inspired the legal norm of managing personal data in EU and incorporated into the GDPR. In practice PbD is far from being the standard in systems engineering in both private and public sector procurement.
We illustrate why privacy needs to be engineered into statistical systems and explain risk when privacy is patched into a solution.
To show how the statistical system could proceed in adopting PbD we bring an example from a high level self assessment of a statistical organization. We walk through future proofing existing services by redesigning an official tourism statistics application that uses mobile location big data.
Finally we envision how statistical organizations can cooperate in supporting each other and learn together how to build trust required in utilising new data sources.
Big data and citizens are inseparable: from smartphones, meters, fridges and cars to internet platforms, the data of most digital technologies is the data of citizens. In addition to raising political and ethical issues of privacy, confidentiality and data protection, the repurposing of big data calls for rethinking relations to citizens in the production of official statistics if they are to be trusted. I argue for relations that involve co-producing data – or ‘citizen data’ – where citizens are engaged in statistical production, from the design of a data production platform to the interpretation and analysis of data. While raising issues such as data quality, I suggest that in a time of ‘alternative facts’, what constitutes legitimate knowledge and expertise are major political sites of contention and struggle and require going beyond defending existing practices towards inventing new ones. In this light, the future of official statistics not only depends on inventing new data sources and methods but also mobilizing the possibilities of digital technologies to establish new relations with citizens.
Since 2013, the Italian National Institute of Statistics (Istat) has been investigating the potential of Big Data sources for Official Statistics. Among such sources, Internet data originated by websites content has been considered as one of the most important to produce information about enterprises. In 2018, Istat started producing experimental statistics on the activities that enterprises carry out through their websites (web ordering, job vacancy advertisement, link to social media, etc.). They are a subset of the statistics currently produced by the “Survey on ICT usage and e-Commerce in Enterprises” and are computed starting from enterprise websites’ contents, acquired by web scraping tools and processed with text mining techniques. A machine learning approach is adopted to estimate models in the subset of enterprises for which the survey and the web sources are both available, with survey data serving as training set for the machine learning task. The content scraped from successfully reached websites is used as input to predict the target values by applying the model fitted in the previous step. The experimental statistics are obtained using two different estimators: (i) a full model based estimator; (ii) an estimator that combines model and survey based estimates. Considering the various domains for which they have been calculated, the three sets of estimates (survey, model and combined) in most cases are not significantly different (i.e. model and combined estimated values lay in the confidence intervals of survey estimates). Simulations have demonstrated that the Mean Square Errors of these new estimates are competitive as compared to those produced in the traditional way.
Internet has been widely recognized as a new data source that can be used either to compile new statistics, or to enhance the traditional ones in several fields of official statistics. Considering that online commerce has a rapid growing share in the overall household’s consumption expenditures behavior broke down by distribution/transaction channel, price statistics is one of the research areas in official statistics which benefits greatly from this new data source. This paper provides a description of the Romanian National Institute of Statistics experience regarding the use of Internet as a data source and an exercise in compiling an experimental consumer price index (CPI) based on Internet data. Aim the pilot project was to investigate whether alternative data collection methods for price statistics can be introduced and enhance the statistical production system in the near future and, most important, it was a great firsthand opportunity to identify methodological challenges which are inherent to Big Data sources from the official statistics point of view. The tool chain is built on top of the traditional methodology used for CPI, enhanced by new features such as simple clustering technique for treating high volatility present in the collected data using a distance-based method for classification similar products.
Following the increasing penetration of the internet, the number of websites that advertise jobs is growing. The European Centre for the Development of Vocational Training (Cedefop) and the ESSnet Big Data have engaged in parallel projects to assess the feasibility of using online job advertisements (OJA) for labour market analysis and job vacancy statistics. After an initial feasibility study finalised in 2016, Cedefop is developing a Pan-EU system providing information on skills demand present in OJA, which will be operational by 2020. The ESSnet has focussed on statistics that can be derived from OJA and entered into a second phase in November 2018 aiming at creating the conditions for a larger scale implementation of the use of OJA in official statistics. This paper builds on experiences gathered in both projects and identifies opportunities and limitations of using OJA for the above-mentioned purposes. In addition, it discusses the feasibility of creating a joint system for processing and analysing OJA data based on discussions that have taken place in the past two years between Cedefop and the ESSnet Big Data on both projects. In this respect, this paper outlines a possible partnership between Cedefop and the European Statistical System to create and manage a unique source of OJA data that would serve multiple uses in the domain of labour market analysis and official statistics. It presents potential types of (statistical) data and variables based on the information contained in OJAs at European, national and regional levels. Data limitations linked to OJAs nature and specificities will be stressed, too. The paper concludes that there is high potential for combining institutional efforts and creating a joint data collection and processing system on OJA and intends to feed a discussion on the feasibility and the implications of creating a European system for OJA, which can serve European and national needs.
The way we eat and what we eat, the way we move and the way we sleep significantly impact the risk of becoming obese. These aspects of behavior decompose into several personal behavioral elements including our food choices, eating place preferences, transportation choices, sleeping periods and duration etc. Most of these elements are highly correlated in a causal way with the conditions of our local urban, social, regulatory and economic environment. To this end, the H2020 project “BigO: Big Data Against Childhood Obesity” (
The 1920s was a decade of great inventions and of substantial productivity growth; people found it difficult to understand why the Great Depression could follow a decade of unprecedented prosperity. Wesley Mitchell and Morris Copeland, who have initiated the flow of funds analysis, urged a better understanding of the circulation of funds between the financial and nonfinancial economy. Since funds, which is the sole currency in the pure credit economy we live today, exist only in the bank’s balance sheets, accounting is a necessity for the virtual currency. Furthermore, the assets and liabilities in the bankers’ accounts mean claims and obligations so that law is another prerequisite for the existence of funds. The present paper is an attempt to detail the historical background of the ‘flow of funds’ analysis tracing back to ancient Rome to clarify the interdependence between law, accounting and economics; and to revive the original idea of Mitchel and Copeland – to understand the interactions between the financial and nonfinancial economy.
When producing anonymised microdata for research, national statistics institutes (NSIs) identify a number of ‘risk scenarios’ of how intruders might seek to attack a confidential dataset. This approach has been criticised for focusing on data protection only without sufficient reference to other aspects of confidentiality management, and for emphasising theoretical possibilities rather than evidence-based attacks.
An alternative ‘user-centred’ approach offers more efficient outcomes and is more in tune with the spirit of data protection legislation, as well as the letter. The user-centred approach has been successfully adopted in controlled research facilities. However, it has not been systematically applied beyond these specialist facilities.
This paper shows how the same approach can be applied to distributed data with limited NSI control. It describes the creation of a scientific use file (SUF) for business microdata, traditionally hard to protect. This case study demonstrates that an alternative perspective can have dramatically different outcomes as compared with established anonymization strategies; in the case study discussed, the alternative approach reduces 100% perturbation of continuous variables to under 1%. The paper also considers the implications for future developments in official statistics, such as administrative data and ‘big data’.
This article proposes some analytical and methodological solutions to study the dynamics of retail business development, based on the results of business tendencies surveys. The following business conditions indicators were developed and tested: Retail Market Indicator (RMI) and Retail Business Potential Indicator (RBPI). The proposed indicators are aimed at quickly identifying current tendencies in the retail business, which, together with the quantitative parameters of the market, increase the scale of representation of the actual and expected phase of economic development of trade and the associated consumer market. The technique was tested to measure the business conditions of Russian organizations in the retail business, over the period 2005–2018. In this study, it was shown that RMI and RBPI are aggregated characteristics of business conditions that can warn about turning points of the business cycle. In addition, based on the decomposition of their dynamics, a tracer of cyclical profiles of indicators, which increases the visualization of industry tendencies, was built and the development of entrepreneurship at different stages of the business cycle was analysed. The results of this study show that the proposed methodology and statistical tools can give a significant contribution to the improvement of the existing techniques on industrial processes’ monitoring.
The Sustainable Development Goals (SDG) indicator framework represents a major challenge and a unique opportunity for the advancement of the global statistical system, both in terms of methodological development and governance. Over the past three years, the Inter-Agency and Expert Group on SDG indicators (IAEG-SDG) has gradually developed a number of documents providing criteria and guidelines for regulating data flows between countries and custodian agencies needed to inform the global SDG reporting process. The validation of methods and data for SDG indicators, while apparently consisting of two completely separate matters, have been closely linked in the SDG process. When validating country data, National Statistics Offices (NSOs) are effectively also certifying the specific methodology used by the custodian agency for the compilation of the indicator, in particular the data source used and the adjustments made to harmonize national definitions and classifications. This article highlights some of the main challenges in the practical implementation of the guidelines on data flows, identifies areas in need of further guidance from the IAEG-SDG and provides some proposals aimed at improving the global SDG reporting process.