
Research article
Select search scope: search across all journals or within the current journal



Under ever-increasing pressure to provide more with less, to justify budgets and to earn public trust, official statistics has long been concerned with how to better prove and communicate its value. Being statisticians, our inclination has been to express the value of our offer in quantitative terms. But an ONS-led Task Force under the Conference of European Statisticians (CES) argued that before we can quantify ‘the value of official statistics’ we need to understand what this really means. This entails first articulating our own central goals as providers of a public good and then working outwards from these goals to formulate the means of fulfilling them. Only then can we start to define measurable indicators of achievement to assess how far we are creating this intended value. This is the reverse of the process often followed, which starts out by identifying already-available indicators and tries to determine the aspects of value of which they are indicative. Future international work should focus on developing tools for better understanding the pathways from goals to value indicators; sharing experiences of efforts to prove and improve the value of official statistics; and developing a core set of measures using the methods outlined in the article.
This paper offers an initial overview of different ideas about the value of Official Statistics systems. For the purposes of this paper, an Official Statistics ‘system’ is taken to cover the set of organisations (within a country) that are involved in the production, communication, use and governance of Official Statistics. The paper seeks to analyse the stated ambitions for, or the claims about, these systems that are contained within a sample of formal corporate documents mainly produced by different national statistical organisations.
These sources offer a range of diverse ideas about different types of value that societies may secure from having well-functioning Official Statistics systems. There are some foundational ideas of value that are often referenced – including the ambition that good Official Statistics will enable good (or better) decision-making by governments and others, which in turn will generate positive outcomes for society.
The analysis also flags a range of ambitions for wider outcomes that might be secured by well-functioning Official Statistics systems – for example, outcomes for citizens (enabling them to be better informed, represented and empowered), outcomes for governments (contributing to a better more effective governmental process) or outcomes in terms of having better informed public debate.
Collectively these concepts could inform any wider framework developed to communicate the potential value of Official Statistics systems, complementing the more specific expressions of value that might emerge directly from the views and judgements of users.
Official statistics are widely considered to be public goods, however this paper explores a higher aspiration: that they also serve the public good. To achieve this goal, and provide value to societies worldwide, there is a need for discussion around what it truly means for statistics to serve the public good. This paper shares initial perspectives on the matter from the United Kingdom Office for Statistics Regulation (OSR) before demonstrating how serving the public good fits with customer-centric perspectives on value, and calling for interested parties to join this discussion so that we may work together in service of statistics for a global good.
The main purpose of this article is to show the quantitative relationship between political regimes and the quality of the national statistical systems. The data exploratory analysis, usually treated from a qualitative point of view, shows a strong correlation between democracy and official statistics, a thesis confirmed in all continents. The most democratic countries are the ones with the best statistical performance. In fact, among the top 10 democracies, five are also in the top 10 for statistical performance. The correlation between democracy and statistical performance is about 70 percent, although over the years 2016–2022, it has been slightly decreasing. The country’s statistical performance is affected by political regime. The main indicators employed for this analysis are the Democracy Index by The Economist, and the Statistical Performance Indicators by the World Bank. The use of these global indicators, encompassing an entire range of years and countries, is unusual in statistical analysis.
Does the current dialogue on the development of statistical systems provide adequate scope for transforming official statistics to deliver their social role? Can statistical systems, as currently defined, provide opportunities to people and non-state institutions to influence “what” statistics and “how” should be produced and used? This paper provides a sociological framework to investigate these questions within a broader understanding of the social functions of official statistics as part of public statistics required for a democratic society.
In recent years, textual analysis and embedding spaces have become essential and complementary tools for sentiment analysis in National Statistics Institutes’ research, owing to their ability to summarize discussed topics effectively. Istat has developed an innovative tool, wordembox, which allows external users to explore the outputs of popular word embedding algorithms, such as Word2Vec and FastText. This tool enriches the analysis with a novel graph functionality, enabling users to discover clusters of words and facilitating implicit topic modeling.
This article focuses on Social Mood on Economy (SME) posts over a period in which the index recorded a strong downward trend: the first month of the Russia-Ukraine conflict at the beginning of 2022. We compare findings from wordembox with standard topic modeling techniques, including Bayesian Latent Dirichlet Allocation (LDA), Top2Vec, and BERTopic, recent methods that extract clusters from word embedding spaces. These techniques show coherent results, and their combined use in textual analysis may create a synergy that enhances the informative content of synthetic indexes such as ‘Social Mood on Economy Index (SMEI)’.
There is a growing demand for statistics to better understand the globalisation that is accelerating due to the removal of barriers in international trade and to technological progress. Key players in globalisation are the multinational enterprise (MNE) groups that have increased in number and complexity and need to be properly represented by macroeconomic and business statistics. To deal with this need, the European Union Member States, the European Free Trade Association (EFTA) countries and Eurostat have collaborated to create the EuroGroups Register (EGR). This paper explores the use of web intelligence to improve the accuracy and completeness of the EGR, which makes use of tools for extracting and exploring information from the World Wide Web. Additionally, it presents a methodology to assess the quality of the information retrieved from the web, based on an
The paper draws attention to the use of Symbolic Data Analysis (SDA) in the field of Official Statistics. It is composed of three sections presenting three pilot techniques in the field of SDA. The three contributions range from a technique based on the notion of exactly unified summaries for the creation of symbolic objects, a model-based approach for interval data as an innovative parametric strategy in this context, and measures of similarity defined between a class and a collection of classes based on the frequency of the categories which characterize them.
The paper shows the effectiveness of the proposed approaches as prototypes of numerous techniques developed within the SDA framework and opens to possible further developments.
A recent application in machine learning has introduced a novel approach, complemented by big data sources, aimed at providing precise estimates for small geographical areas. This method employs a dual strategy: (a) hybrid estimation, involving the integration of big data sources with imputed values derived from
This paper enhances the comparative analysis by contrasting the CKNN method with a hierarchical Bayes method using the logit-normal model (LN) relevant for binary data. Broadly speaking, the LN method can be viewed as the Bayesian equivalent of Battese-Harter-Fuller (BHF) method, which incorporates unit-level covariates. Our results demonstrate the CKNN method’s superiority over the LN method. However, the application of hybrid estimation to the LN method significantly diminishes this superiority. Although CKNN estimates maintain better precision, they are not as accurate as the estimates from the hybridized LN method.
The rapid technological changes have revolutionised how we function, including how we search for work and what skills we need to be equipped with to perform the tasks at the workplace. As employers more often recruit using online job advertisements, their content becomes a natural source of information for analytical purposes on the skills demanded in the labour market, especially for analysing emerging skills like digital. There are still some challenges with the extraction of information from online content. However, the extraction improvements go hand in hand with new technological developments like natural language processing techniques. This article presents the experimental method of updating the classification of digital skills to keep it up to date for information extraction applied to online job advertisements. The evaluation proved this method successfully identified terms related to programming skills but failed to identify terms associated with artificial intelligence sufficiently. The latter is related to the fact that the AI field is among the fastest developing areas of technology advancement, and new terms (e.g. Chatgpt) always appear.
Gross Domestic Product (GDP) stands as a pivotal indicator, offering strategic insights into economic dynamics. Recent technological advancements, particularly in real-time information dissemination through online economic news platforms, provide an accessible and alternative data source for analyzing GDP movements. This study employs online news classification to identify patterns in the movement and growth rate of Indonesia’s GDP. Utilizing a web scraping technique, we collected data for analysis. The classification models employed include transfer learning from pre-trained language model transformers, with classical machine learning methods serving as baseline models. The results indicate superior performance by the pre-trained language model transformers, achieving the highest accuracy of 0.8880 and 0.7899. In comparison, hyperparameter-tuned classical machine learning models also demonstrated commendable results, with the best accuracy reaching 0.845 and 0.7811. This research underscores the efficacy of leveraging online news classification, particularly through advanced language models. The findings contribute to a nuanced understanding of economic dynamics, aligning with the contemporary landscape of information accessibility and technological progress.
Many censuses and surveys in low- and middle-income countries ask questions about deaths in the household to fill the evidence gap about mortality. This study undertakes the first published systematic assessment of the completeness and quality of these data. For 82 censuses from 56 countries and 26 surveys from 21 countries since 2000 we calculated completeness of household death reporting using deaths estimated by the United Nations World Population Prospects (UN WPP) and Global Burden of Disease (GBD) as the denominator. The median completeness of reported household deaths in censuses was 89% (inter-quartile range (IQR) 66–102%) and surveys 96% (IQR 80–124%). Completeness was similar for males and females and substantially lower where date of death was asked (census median 73%, IQR 53–91%) than not asked (census median 93%; IQR 74–110%); these differences remained after controlling for other covariates in a linear regression. The ratio of reported household to estimated deaths was higher in younger ages but age-invariant where date of death was asked. In conclusion, household death data in censuses and surveys have major completeness and quality issues. Where date of death was not asked, there appears to be considerable reporting of deaths that occurred outside of the reference period.
While national registry systems are evolving worldwide and, in some cases, replacing reliance on censuses, in countries where well-established population registers are lacking, the population and housing census remains the primary source of detailed data on the number of people, their spatial distribution, age and gender structure, living conditions, and other key socio-economic characteristics. The quality of the census findings is crucial for several reasons, including building public trust in the national statistical system. In many developing countries, conducting a Post-Enumeration Survey appears to be the only feasible way to evaluate the census results. Indeed, the lack or incompleteness of reliable demographic data from alternative sources precludes the use of other methods. This paper discusses some aspects of the feasibility of a Post-Enumeration Survey in Ethiopia. In particular, the paper reports on the main critical issues that emerged from the pilot surveys carried out in the framework of a cooperation project – funded by the Italian Agency for Development Cooperation – aimed at providing methodological support and technical assistance for the preparation of the 4th Ethiopian Population and Housing Census.
Since January 2022, the Regulation on European Business Statistics (EU 2019/2152) requires EU Member States to compulsorily share microdata on intra-EU exports. Establishing intra-EU export Micro-Data Exchange (MDE) provides National Statistical Institutes with a new data source to compile intra-EU import statistics. The availability of MDE tackles two key challenges: diminishing the overall response burden on data providers and meeting user expectations regarding the quality of the produced statistics. However, transitioning to a data production system based on MDE data requires the assessment of the coherence and comparability between MDE and National import data.
To identify asymmetries between the two data sources, Istat developed an innovative application designed to foster cooperation among Member States. The tool was developed using the Shiny package in R. The implemented solution allows users to perform exploratory analysis, systematic error detection, and selective editing. The most relevant asymmetries are identified through relative contribution and the asymmetry suspicion indices assessed by
Sharing the open tool within the European Statistical System enhances interoperability, promotes method harmonization, and encourages the adoption of official statistical standards.
Deciphering energy efficiency is a critical component for the sustainability and energy policies of the European Union (EU) and its Member States. Decomposition analysis is a key method that helps distinguish real energy efficiency gains from other external factors. This article presents a decomposition analysis method using official EU statistics, its results, and associated limitations. Although separating structural changes and activity levels from energy efficiency poses a challenge, an adapted Logarithmic Mean Divisia Index (LMDI) method was utilised to isolate and highlight energy efficiency. This article outlines this method and how the European official statistics were exploited. Firstly, it examines the factors affecting energy consumption in various sectors within EU-27. Secondly, it examines the factors affecting energy consumption in various sectors. In doing so, it also addresses potential obstacles in data collection, and presents improvements to the LMDI analysis. The findings in this study make a substantial contribution to the fields of national statistics, methodological applications, and energy data analysis in the context of the EU’s energy policies.
Gross domestic product (GDP) is unquestionably one of the most influential statistical indicators in history. It is more than a statistic – it not only measures the global economy but defines it. But from the outset there have been criticisms of GDP. Today there are a growing number of commentators arguing that GDP has outlived its usefulness. Their criticisms can be broadly categorized into three classes. The first are measurement problems within the existing framework arising from changes in the economy and society – most notably globalization and digitalization. The second set of criticisms deal with the limits of the SNA framework itself and are sometimes described by the catchall “Beyond GDP” and center on questions as to whether the SNA can or should measure well-being and sustainability. The third is that the construction of GDP promotes a ‘growth-at-all-costs’ ideology which works against environmental and social reforms.
This paper summarizes the origins of the SNA and GDP and some of the crucial events and thinking that helped shape its design. The most important criticisms and challenges that will shape the future development of the SNA are also outlined, in particular: globalization, digitalization, well-being and sustainability. As both well-being and sustainability go well beyond traditional measurements of the economy, the paper discusses whether it is possible to address at least some aspects of these issues within the SNA, either in the ‘core’ sequence of economic accounts, or through a broadened set of accounts. The paper concludes with an overview of the 2025 SNA update and new work beginning at the UN to encourage member states to move beyond GDP.
The phenomenon of Business-to-Government (B2G) data sharing represents a growing trend, especially in latest years. In fact, research has shown how privately held data could have a huge potential when used to tackle societal policy issues. B2G data sharing initiatives can be employed in different situations: from emergencies to the construction of official statistics and the use in research, just to name a few. In all these circumstances, the quality level required for the data may be different, as different principles could prevail upon others (e.g., timeliness in the case of emergencies is a key parameter). This heterogeneity in possible use-cases motivates the present work. In fact, our objective is to understand and classify the different contexts in which B2G data sharing may happen. The idea is to create a taxonomy of B2G data sharing initiatives, in which we identify all the different instances where B2G data sharing may occur. Afterwards we add as attributes some identified quality principles that characterise the different B2G data sharing situations. The work aims at providing further information that can help clarify specificities and requirements of B2G data sharing in order to enable relevant data flows and make them more dynamic.
The global data ecosystem is changing rapidly. New demands are increasingly being placed on National Statistical Offices (NSOs) worldwide to collect data to track a growing array of indicators. However, many NSOs lack the capacity to collect frequent, representative and high-quality data on even core metrics of national progress, such as food security. In response to this, there has been a growing number of partnerships between NSOs, international organisations such as the United Nations, and private sector organisations to address the data gap, which was only accelerated by the COVID-19 pandemic. Using Gallup, the global research and analytics firm, as an example, this paper highlights a number of areas where the private sector can provide value to the realm of official statistics. By adhering to globally recognised statistical protocols with a firm commitment to the principles of rigour, transparency, and respondent confidentiality, organisations such as Gallup play an important role in supporting the collection of official statistics. They can also bridge key data gaps related to the most pressing challenges of our time, and drive accountability on key issues of national and global development.
The datification processes have driven National Statistical Institutes (NSIs) to explore new data sources with the goal of enhancing the quality of statistical information. The creation of innovative statistical products, known as Trusted Smart Statistics (TSS), necessitates new production processes that integrate with traditional methods, maintaining levels of relevance, quality, and trust.
To facilitate this transition in statistical production, NSIs must embark on change management initiatives involving all organizational dimensions engaged in statistical processes. In the scientific debate, sustainability has assumed a broad connotation. This paper delves into the organizational aspect, viewing sustainability as a paradigm that encompasses all organizational areas and their interactions crucial to supporting the shift to new statistical production models, without compromising traditional practices.
The paper explores the path taken by Istat in this direction, the organizational solutions implemented, the benefits realized, and investments in research and innovation initiated years ago on new data sources, aligning with developments in the European statistical community. Istat has been at the forefront of big data experimentation in Europe for over a decade, implementing a modernization program and establishing research infrastructures. Istat took a significant step by founding the TSS Center to guide organizational adaptation at technological, methodological, legal, communication, and human resources levels toward the new production system.