Abstract
German official statistics publish statistics on personal insolvency. These statistics have been recently enhanced using web scraping to extract additional information from a public website on which the insolvency announcements are published. The currently scraped data is used for quality assurance and to derive an early indicator of personal insolvency. This paper provides novel methodological analyses for the same administrative database and presents further opportunities to improve the current official statistics regarding detail and timeliness using web scraping and text mining. These newly derived statistics inform on several aspects regarding personal insolvency’s demographic and spatial distribution.
Introduction
In official statistics, the need for new data sources and methods to improve cost efficiency and timeliness of statistical production has been addressed for over a decade. In 2013, the Directors General of the National Statistical Institutes of the EU signed the Scheveningen memorandum, declaring that “the demand for timely and cost-efficient production of high-quality statistical data increases” [1]. Since then, there have been several initiatives to improve official statistics using new and mostly non-probability-based data sources or data science techniques [2, 3, 4, 5].
In this context, web scraping has been identified as a promising method in official statistics, and several national statistical institutes have adopted research projects using web scraping to utilize the vast public data available on the internet. Meanwhile, there are several implementations and use cases for official statistics: For example, agritourism [6], job vacancies [7], sustainability reporting [8], consumer price indices [9, 10], detection of innovative companies [11], or real-time house prices during the COVID-19 pandemic [12]. For recent overviews on web scraping in official statistics, we refer to [13, 14]. Potential limitations of web scraping to be used in official statistics are discussed by [15]. The interest of official statistics in web scraping is mainly due to the automated data collection process. Thereby, it is expected that the response burden can be reduced, data can be collected faster, and thus statistics can be produced and published more quickly. Furthermore, new data sources can be gathered, resulting in opportunities for new official statistics publications, and ultimately, costs can potentially be reduced.
This paper analyzes web-scraped data from an online database that publishes administrative records on personal and business insolvencies. The database is publicly available and is used by the German Federal Statistical Office (Destatis) as an early indicator for insolvencies [14] and as a form of insolvency-nowcast during the COVID-19 pandemic [16]. However, the database still contains considerably more information, and the full potential of the database does not seem to have been exploited yet. This applies, in particular, to information on the demographic and spatial characteristics of the units contained in the database. This information is retrieved from large text fields stored in the database using text-mining techniques for the present study. The first evaluation and analysis of this information were presented by [17], with more details in [18].
The database in question contains all business and personal insolvency cases in Germany. More specifically, it contains all announcements on insolvency proceedings for the last two weeks. The database is updated regularly, which must be considered when scraping the data. For this study, we scraped one year from the database and focused on personal insolvency. Personal insolvency is a simplified procedure for handling the insolvency of a natural person (private person). It is intended to equally satisfy the creditors of an insolvent debtor on a pro-rata basis. Declarations of insolvencies have to be made public in Germany (§9 Public announcement) [19]. After the respective court decision, the declared cases are stored in a database and are published on the website,1 operated by the Ministry of Justice of the federal state of North Rhine-Westphalia [14]. Certain businesses are even obligated to apply for this procedure (15a Obligation to file applications for legal entities and companies without legal status) [19]. For further details on procedures and further legal aspects concerning personal insolvency in Germany, we refer to [16].
Most non-administrative online data sources concerning economic indicators are based on data gathered by private companies or other non-administrative data sources. However, the database scraped for this study exclusively publishes administrative records. This way, in contrast to many other studies, we can use an online database with the quality and completeness of administrative data. As a result, problems associated with web scraping applications, such as coverage or selectivity [20], are not of concern in this study.
This paper’s major contribution is combining web scraping and text mining to extract information from an existing administrative online database that has not been analyzed and published. We focus on demographic and spatial information to provide results of the corresponding distributions. Furthermore, we show that the database contains information to produce high-frequency time-series data.
The paper is organized as follows: First, we inform of the societal relevance of personal insolvency and previous research (Section 2). Section 3 describes the developed web scraper and some technicalities. The scraped data in its content, the analytical strategy, and other data sources used are described in Section 4. The results are presented in Section 5. Finally, the results, strengths, and limitations are discussed in Section 6. Conclusions are presented in Section 7.
Research background
In the last decades, strong increases in personal insolvency rates have been reported by several researchers [21, 22, 23, 24]. This rise in insolvencies has been accompanied by increased public interest and research, as this topic is closely related to issues of social inequality and economic stability. Apart from the USA, Australia, and Canada [25], Germany and Great Britain are regularly mentioned as countries with high insolvency rates for the European context [26, 23].
According to [21, pp. 2–3], several branches of research in the area of personal insolvencies attempt to answer various overlapping questions. For example, some studies focus on the impact of an insolvency proceeding on an individual’s future credit score (e.g., [27, 28]), while others compare the implications of varying insolvency laws (e.g., [29, 30]). Most empirical results are focused on the US context, given the limited availability of data on insolvencies in Europe [21, 23]. As argued by [23], this is also pronounced regarding the spatial distribution of insolvencies; however, empirical research on this topic outside of the United States has been limited due to a lack of data.
Various approaches exist in the research literature to explain insolvency rates and the associated increases in recent decades. One of these approaches is the “traditional model of consumer bankruptcy”, which is, for example, presented and discussed by [31]. According to this view, insolvency proceedings are utilized by individuals when they are in financial distress and cannot otherwise resolve the situation [31, 23]. However, it is emphasized that the traditional approach, which has long provided a valuable explanatory contribution to explaining bankruptcies between phases of economic decline and improvements, no longer seems to describe recent decades’ changes well [31, 32]. Thus, more research is required to respond to these recent trends with adequate explanatory approaches.
Several aspects are discussed as contributing factors, which usually affect individuals externally and can include, for example, (severe) health problems or unemployment, which are also associated with a loss of income [32, 31, 21]. In a paper by [21], which compares insolvency dynamics between Germany and the United Kingdom, a further distinction is made between macroeconomic factors in which individuals are embedded (e.g., interest rates or credit availability [32]) and microeconomic factors. According to [21], a stronger focus of empirical research is on the latter and generally includes adverse events that an individual may be confronted with (e.g., [24, 33]). It also seems plausible that external shocks, such as the COVID-19 pandemic (which inspired the development of the early indicator of personal insolvencies in Germany [16]) or natural disasters, which have been discussed frequently in recent years, may also structurally affect individuals in certain regions. For example, some regions in Germany experienced severe flooding in 2021 [34].
[21] argues that attitudes and values are also associated with insolvency procedures, and thus, the role of social stigma in personal insolvencies is important (see for example [35]). Therefore, it is likely that the increased insolvency rates in recent decades are associated with a reduction in social stigma [21]. This cultural component is also evident in the fact that in the US, bankruptcy is not necessarily associated with stigma [26]. However, other researchers reported that the available evidence concerning the impact of stigma and insolvencies is rather tentative [23, 36, 37]. Related to attitudes toward insolvency proceedings are the countries’ respective insolvency laws, which shape the extent to which individuals have the option to make this (strategic) decision [32, 38, 36]. In recent decades, European insolvency laws have changed significantly [21, 39]. “Since the early 1990s, many European countries have enacted consumer debt adjustment laws that stand as equivalent to the US consumer bankruptcy in that they give access to (partial) relief of debt” [39, p. 445]. The most recent amendments to the German Insolvency Code (InsO) came into effect in 2022, underscoring the ongoing social relevance and topicality of the issue. For further details concerning the German insolvency law, we refer to [21, 40, 19].
A highly relevant study for our context, which focuses on spatial differences in personal insolvencies in England and Wales, is the work of [32]. The paper investigated whether different economic and sociodemographic factors might help to explain differences in personal insolvency rates. Other studies have also identified the importance of (socio-)demographic factors, with family structure, ethnic background, and education being found relevant [23, 41]. The study by [32] further aligns with more recent work in (spatial) inequality research (e.g., [42]). For example, differences between individuals regarding access to good schools and well-paying jobs [43] also appear at the spatial level, which in turn may be reproduced or further amplified.
[32] used detailed data for the corresponding regions in their study: “Data are for 2006 or 2005/06 and are obtained from NOMIS (demographic, occupational, employment and wage data), The Insolvency Service (insolvency data) and the Land Registry (house prices)” [32, p. 1694]. The study’s descriptive statistics showed significant spatial variation for insolvency rates, indicating several spatial clusters [32, p. 1696]. The authors found evidence that various economic and sociodemographic factors, such as age and occupation, are associated with insolvency rates and that it is important to consider spatial autocorrelation in statistical models accordingly [32].
To the best of our knowledge, there is no comparable research with this level of spatial detail in other European countries, which may be due to the data limitations discussed previously. Thus, our work and the paper by [32] are the only studies to date that analyze spatial differences in insolvency rates in such detail. However, we do not have the same amount of demographic and economic information and thus can only examine some aspects.
Web scraper implementation
Web scraping was implemented in R 4.2 [44] using a Selenium driver to emulate a web browser. The emulated browser pre-fills the search form and submits a request. Appropriate waiting times were implemented to minimize server load from the scraper, as recommended by [45, p. 19, p. 46].
The result page could be extracted easily using our implementation. However, the free-text data for each record containing the data of interest to this study require a database request to be sent. This step can not be emulated straightforwardly. However, using Selenium, each request can be emulated by clicking the required buttons for a single record. Without browser emulation, this would not be possible. The entire free text was extracted for each record and combined with previously scraped results.
Since there is a limit on the maximum number of results that can be displayed using a single request, daily scraping was required, with a lag of seven days for the date scraped. We found some back-filling of insolvency announcements, which could date back a few days. This is most likely due to the data being uploaded manually. Using the scraper daily for the date seven days ago seems to have worked best (see also Section 1 on the temporal availability of the data). In the entire scraped data, some days are missing with no announcements because the website had maintenance days (see Section 6 for a discussion on this).
Data and analyses
The scraped data contains filed cases from one year (week 9 of 2022 to week 9 of 2023). Within this period, 141,124 unique entries related to personal insolvency proceedings have been scraped from the database. Table 1 shows the selected entries. The top three entries are 37% openings of new proceedings, 26% decisions in the proceeding, and 17% decisions in residual debt discharge proceedings. All entries were considered for analysis, and to improve readability, the term ‘personal insolvency’ is used in the following as a synonym for all entries in the database. The following information was extracted from the text fields using text-mining techniques: date of decision, age (derived from date of birth), gender, and geo-location (postal code level). The geographical coordinates of the geo-locations were obtained using the R-library developed by [46]. All spatial shapefiles were obtained from [47].
Categories of personal insolvency proceedings in the database
Categories of personal insolvency proceedings in the database
For analyses, we calculated the incidence rate of personal insolvencies in one year per 100,000 inhabitants. This is the number of insolvency proceedings related to the population size times the considered period. In this case, the period considered is one year, although other intervals are possible (e.g., weekly, monthly, quarterly). The specific and official population sizes (number of inhabitants) per federal state (16 federal states) or per postal code level (8,169 postal codes) were used.
Age distribution (in percentage) of individuals affected by personal insolvency proceedings (shown with bars). Individuals aged 18–100 are shown, binwidth 
Using the location data, we analyzed whether spatial patterns and differences regarding the incidence of personal insolvencies can be identified. The spatial relation is known as spatial autocorrelation. We used Moran’s
with
First, we report results focusing on the distribution of the demographic variables (gender and age) of individuals affected by personal insolvency (Section 5.1). Second, results on spatial distributions of insolvency cases and the spatial distribution of the demographics are presented (Section 5.2). Third, we report the possibility of deriving high-frequent time series data (Section 5.3).
Demographic distributions
In total, there are 70,370 (50%) males and 51,282 (36%) females, and 19,472 (14%) individuals for which the gender could not be determined from the data. Given that gender could not be determined for a considerable number of individuals, this group is considered ‘Unknown’ in the analyses. The missingness is because only the first and last names are given in these text fields, but there is no addressing as Mr., Mrs., debtor, or debtoress (see Section 6 for a discussion on this).
Figure 1 shows the relative distribution of age. There is a steep increase in insolvencies after the legal lower limit of 18 years (German law prevents individuals under 18 from going bankrupt), with a peak in the early thirties. The highest count is found for individuals aged 33 (4,047). There is a plateau in the late thirties to early forties. A decrease in counts until the age of 50–60 years is observed, with another plateau for these age groups. There is no noticeable increase in insolvencies for retirees (ages 63 to 67, depending on the birth cohort), as insolvencies steadily decrease with age. On average, the individuals are aged 44.5 years, and the oldest individual is 111 years old (not shown in the Figure). The age could not be determined for about 7% of the data. For comparison, data from a German general population survey, the ‘ALLBUS’ [50], is also shown. There are two distinct differences. In relative terms, there are more younger individuals in the insolvency data than in the general population. On the other hand, there are fewer older individuals in the dataset compared to the proportions in the general population.
Age distribution (in percentage) of individuals affected by personal insolvency proceedings split by gender (shown with bars). Individuals aged 18–100 are shown, binwidth 
The incidence of personal insolvency proceedings. The incidence determines the color gradient. From light-gray (indicating lower values) to black (indicating higher values). The geographical borders of the postal-code areas (panel b) are not shown for visualization purposes.
The incidence of personal insolvency proceedings by age and gender. The incidence determines the color gradient. From light-gray (indicating lower values) to black (indicating higher values).
Considering the age distribution by gender, there are no noticeable differences in the percentages between the three groups (see Fig. 2). No gender has an age group containing more than 3.5% of the observations. The distribution of the group ‘Unknown’ does not differ noticeably from ‘Male’ or ‘Female’, suggesting that there is no substantially different age distribution between these groups. For comparison, the German general population data [50] is shown as well, except for the group ‘Unknown’. For both genders, the same pattern is observed; in relative terms, more younger individuals are in the insolvency data than in the general population, and vice versa, fewer older individuals are in the insolvency data than in the general population.
Figure 3 shows the spatial distribution of personal insolvencies on the federal-state (a) and postal-code level (b). The color gradient indicates the incidence in each federal state or postal-code area per 100,000 inhabitants. On the federal-state level, the maximum of personal insolvencies is found in Bremen (274 per 100,000 inhabitants) and the minimum in Bavaria (92 per 10,000 inhabitants). There are 189 personal insolvencies per 100,000 inhabitants averaged over all federal states. Both spatial resolutions show a north-south disparity. However, the incidence values on the postal-code level were trimmed to 500 for visualization purposes. Otherwise, some large outliers would have too large an impact on the color gradient. There are 101 postal codes with an incidence larger than 500. In the federal state of Saxony, an unusual spatial pattern is evident. There are 234 out of the 403 postal-code areas in Saxony, with an incidence of 0 (which is a valid value in the absence of the attribute), while ten postal-code areas have an incidence of 500 or larger. For example, one postal code in Saxony has an incidence of about 6,900. In general, low incidences and less variation in incidences are observed for this federal state. These spatial patterns cause the geographical boundaries to be visible. This result is not observed in any other federal state (see Section 6 for a discussion on this).
Figure 4 shows the spatial distribution of the age and gender of individuals affected by personal insolvency proceedings per federal state. Evidence for a north-south disparity also seems to exist, considering the median age. In the mid-northern part, the median age is 41, while the median age is about 45 in the southern states. However, Berlin seems to be an exception, as the highest median age is found here. Considering the spatial distribution of gender, the ratio of male individuals is used for the color gradient. The northern- and southern parts seem to have larger ratios of male individuals, while the states in the mid-part show lower ratios of male individuals.
To test whether differences in the observed spatial patterns can be identified, Moran’s
Time series
Time series data on annual/monthly relative differences of insolvency proceedings filed are published [16]. However, the database allows for even more detailed and high-frequency statistics. Here we show an example, considering the weekly counts of filed declarations of personal insolvencies (see Fig. 5). The period from week 10 in 2022 to week 8 in 2023 is shown. The weeks (09-2022 and 09-2023) on edge are not shown, as the counts differ considerably (see Section 6 for a discussion on this). Even though both the counts and the trend indicate an increase, we consider these results preliminary and experimental. Stronger conclusions about development and trends for this high-frequent series could be drawn upon the existence of a series for several years. These time series can, in principle, be created for all categories listed in Table 1. For example, such time series data can be used to produce preliminary but timely predictions of the present (nowcasts) for sample surveys that usually provide yearly or quarterly estimates. Such results could be used as early warning indicators for policy-makers. Furthermore, this time series data can be used as a target or auxiliary series in structural time series models. As a target series, it could be used to study the effects of targeted policy interventions on insolvency counts or as an auxiliary series to study its effects on the series of unemployment rates. However, when using such time series data, one should be aware that counts of insolvency proceedings must be considered as a lagging indicator (for details, see [23]).
Weekly counts of filed declarations of personal insolvencies. Black thin line shows weekly counts, the black bold line is a smoother using ‘loess’ (span 
The main goal of this paper was to demonstrate the possibility of producing detailed and potentially more timely official statistics on spatial and demographic distributions of personal insolvency using an existing and, for the demonstrated purposes, untapped large administrative database. The fact that personal insolvency is a current social problem has been described in detail in Section 2. The obtained statistics are based on combining web scraping and text-mining techniques. According to the main findings of this study, there is evidence that there are certain age groups, for example, younger individuals in their early thirties, where personal insolvency is relevant. Furthermore, there is evidence for spatial clustering of personal insolvencies, which indicates regional disparities. These results are consistent with previous results on factors associated with personal insolvency (e.g., [32, 42]). Regarding the spatial distribution of age and gender our findings are mixed, but only one year is considered in this study, and the results reported should be considered preliminary, indicative, and experimental statistics. However, this is the first study that has provided such a level of detail for Germany. Thus, the main goal of this paper has been successfully achieved. An advantage of the data used in this study is that the data are a presumably almost complete annual register. Thus, official statistics could be produced much more straightforwardly as with most so-called ‘new data sources’, which are mainly based on non-probability samples [52]. However, this paper showed that the potential integration in producing official statistics is not straightforward, even in such a fortunate situation. Therefore, we will elaborate on the limitations and give recommendations for future research.
First, the database is based upon manual uploading processes and is based on the accuracy of the information uploaded into the database. The latter may give rise to problems, which are addressed in several points below.
Second, the text-mining rules applied to derive gender and age did not yield results in all cases. For gender, 14%, and age, 7% of the entries had missing data. Although the results in this paper suggest that there are probably no differences across groups, further research is needed here. For example, it needs to be clarified whether the text mining rules are inappropriate, whether there are input errors, or if the input is missing. Furthermore, it must be clarified whether the missing data possibly occur in different categories (see Table 1).
Third, considering the results on the postal-code level (Fig. 3, panel (b)), there is a suspicious result for one federal state (Saxony). The spatial pattern deviates from the remaining country. It needs to be clarified whether this is due to errors in the database, inappropriate text mining rules, or whether this corresponds to empirical reality.
Fourth, it should be noted that the decision date contained in the text fields should be used when producing time series. Using the upload date can deviate considerably from this (see Section 3), leading to time clusters in the data (the counts then depend on the size of the upload batches) which ultimately result in incorrect time series data. We also found that the scraped data’s first and last weeks contained considerably fewer counts. Possible reasons for this may be upload timings, but we cannot conclusively clarify. Furthermore, the text field for the date contained irregular dates (entries that were too far in the past or future).
We do not consider these points severe limitations to the study and analysis but logistical problems to be addressed for potential implementation, for example, the standardization of input data or the development of an API for this database with exclusive access to official statistics. This last point is crucial, as the purpose of the website and insolvency announcements, in general, is not the public availability of data to create a database but primarily serves information purposes described in the German Insolvency Code (InsO). The main goal of insolvency proceedings is to enable the debtor to free themselves from their remaining liabilities and, at the same time, satisfy their creditors [19]. Although these publications should create transparency, they are also subject to certain deletion periods and thus aspects of data protection, so a publicly accessible API that enables the mass storage of insolvency entries is neither intended nor desirable. For official statistics, processes suitable from a data protection perspective would have to be developed across all authorities and then consistently applied in the production process.
One last point that we would like to make an object of our discussion is that it would be valuable if the general availability of data on insolvencies in the European context were to be improved so that more in-depth analyses in individual countries were possible, which would then also allow for cross-country comparisons. The data available in Europe is currently limited (as described in Section 2), and German official statistics could serve as an example with this database and the statistical products that are based on it. As several changes in insolvency laws have been adopted over the past 30 years in Europe, the situation concerning data is still not optimal, as described by [21]. This seems even more important from a policy perspective and considering the argument that insolvency rates may reflect economic difficulties in certain regions, which in turn may be related to other social problems [32, 23]. Describing this in detail could provide a valuable foundation for decision-making processes.
Conclusion
This paper demonstrates opportunities for German official statistics to publish new statistics on personal insolvency’s demographic and spatial distribution. A potential implementation in official statistics seems feasible as this database is already used for quality assurance and to derive an early indicator of personal insolvency. However, the feasibility of a future production model depends entirely on the logistics. An exclusive API for official statistics and cooperation between authorities would be the basis. This basis could result in a continuous and standardized data stream processed using standardized routines, resulting in official statistics presented in this paper. Future research should focus on the robustness of current analytical strategies and consider further options for analysis, for example, investigating whether this data is suitable for forecasting projects or the analysis of business insolvencies and how this information could be used for socio-political questions.
Footnotes
See
Acknowledgments
The views expressed in this paper are those of the authors and do not necessarily reflect the policies of their affiliations.
We would like to thank the reviewers for taking the time and effort. We appreciate all valuable comments and suggestions, which improved the quality of the manuscript.
