Abstract
In a globalised and interconnected world, enterprise groups play an increasingly important role in the economy. At the same time, it is becoming more difficult for official statistics to map these structures correctly to be able to analyse them for the various statistical areas. Enterprise groups have complex data structures that often come from different data sources of varying quality and need to be compared at different points in time. Comparing these complex and growing entities is not trivial and has therefore rarely been done in the past across countries and sources. To fill this gap, inspired by other similarity metrics, a new similarity metric has been developed in the German statistical business register: RUMS –
Keywords
Introduction
The statistical business register is the backbone of official business statistics and covers all the different units used in business statistics. 1 One of these unit types is the enterprise group, which is becoming increasingly important in the economy. Enterprise groups are often multinational, with structures spread across many countries.
An enterprise group is defined as a set of more than one legal unit (LEU) linked by control relationships. A LEU is controlled by another LEU if more than 50% of the voting rights are held by that other LEU. Within an enterprise group, every LEU is controlled by another LEU (parent). There is one LEU on top of the control chain. This is called the global group head (GGH). The GGH and the controlled LEUs together form the enterprise group. 2
An enterprise group in a statistical business register usually has specific properties, e.g. turnover, employees, main economic activity (e.g. NACE-Code), identification number and so on. Within an enterprise group some LEUs act in specific roles like GGH or global or regional decision centre (head office).
In both data production and assessment processes, e.g. for statistical publications, it is essential to compare enterprise group data either between different registers in order to evaluate the registers’ different data sources, or at different points in time within the same register in order to evaluate how the structures of enterprise groups have changed over time.
Business registers as a population contain a large number of enterprise groups (in Germany, around 290 000 in reference year 2022) that need to be compared. When trying to find the same enterprise group, a lot of the enterprise groups compared do not have exactly the same structures. This can be caused by changes in the group structure over time or by quality differences between data sources. To be able to compare different data sources or different points in time, rules must be defined to identify the same or diverging enterprise groups. There exist earlier straightforward approaches that used simple hypotheses like the identity of the GGH to compare enterprise groups. 3 However, a deviation of the GGH does not always indicate a different enterprise group. In our point of view other properties are more relevant and should be considered when comparing two enterprise groups.
To illustrate, Figure 1 shows two enterprise group structures. In enterprise group B, the structure of LEUs 1, 2, 3, 5 and 6 is the same as in enterprise group A. However, the LEUs 4, 7 and 8 are missing compared to enterprise group A and there is LEU 9 that does not appear in the group structure of enterprise group A. How similar are the two enterprise groups?

Two enterprise group structures.
Due to improved data storage for the statistical unit enterprise group in the German Statistical Business Register, it became necessary to define the consistency rules for enterprise groups more precisely. For this purpose, RUMS was developed and used for the first time. The new similarity metric RUMS quantifies the similarity of enterprise groups and makes them comparable. This makes it possible to compare Enterprise Groups within the same business register over time and to compare Enterprise Groups between different business registers at the same time. The idea of a similarity metric, its requirements and the formula for calculating RUMS values are explained in chapter two. Chapter three presents two use cases of RUMS, one within the German national statistical business register (NSBR) and one comparing the German NSBR with an international business register. Subsequently, chapter four presents RUMSA, a meta measure that can be used to compare all enterprise groups across two sources or points in time. The fifth and final chapter summarises the main points of the paper and gives an outlook on possible further use cases and developments of RUMS and RUMSA.
Similarity and distance metrics attempt to compare two objects and quantify the similarity or distance between them. In the literature, there are measures for comparing character strings or entire text documents, 4 but also for more complex objects such as faces, plants or genes. 5
A distance metric d(x,y) measures a distance between two objects, either in terms of physical distance i.e. how far you have to go from x to y, or how much you have to change an object x to get y. 6 The distance measure is 0 for two identical objects, and the greater the difference, the greater the calculated distance. Distance metrics describe the dissimilarity between two objects and can be easily converted into a similarity metric s(x,y). A non-normalised similarity distance metric has no maximum value.
A normalised similarity metric s(x,y) is always between 0 and 1 such that 0 ≤ s(x,y) ≤ 1 for all objects x and y. This is done by taking the intersection of the objects x and y and dividing it by the total number of considered properties. The higher the value of the similarity measure s(x,y), the more similarities the objects x and y have. If the objects differ in all considered properties, the similarity s(x,y) = 0. If the considered properties of objects x and y are exactly identical, s(x,y) = 1.
RUMS, a similarity metric for comparing enterprise groups, is designed to fulfil the following expectations.
If two enterprise groups are completely different, i.e. have no overlap in the characteristics considered, then RUMS should be 0. If two enterprise groups partially overlap in the properties considered, i.e. are similar to each other, RUMS should be between 0 and 1. If two enterprise groups are completely identical in terms of the properties considered, then RUMS should be 1. The higher the similarity between enterprise group A and enterprise group B, the higher RUMS should be. This means that if enterprise groups A and B are more similar than A and C, RUMS(A,B) > RUMS(A,C) must always be true. RUMS should be flexible enough to expand and reduce the properties considered. It should be possible to flexibly weight the considered properties in RUMS, depending on the expected quality and importance of the properties in the populations being compared.
Requirements e and f are intended to allow the widest possible implementation for a variety of use cases in official statistics.
RUMS uses quantitative characteristics available in the German NSBR. To calculate RUMS between enterprise groups A and B, the number of overlapping LEUs L(A∩B), the number of overlapping employees E(A∩B) and the turnover of the overlapping LEUs T(A∩B) are related to the total number of LEUs, employees and turnover in enterprise groups A and B respectively. The general formula is:
Equation 1
If one of the denominators is 0, then there is one of the enterprise groups with no employees or turnover. To handle these cases where the denominator is zero, we define the entire fraction as 1 by convention. This way, such constellations remain calculable and do not result in this property not being considered or other properties having to be weighted higher as a result. This would distort the comparability of RUMS values. The following example (Figure 2) shows the calculation of RUMS for the comparison of enterprise group X with Y and with Z. For the example, the parameters for weighting the properties and the two populations are equated, meaning a, b, c, d, e, and f are all

Three enterprise groups with overlapping LEUs.
Equation 2
The three properties considered (LEUs, employees and turnover) and their intersections in RUMS formula are each set in relation to the total of both enterprise groups and are thus considered in two sub-indicators, so that additional LEUs on either side reduce the similarity. This allows one of the two populations to be weighted more strongly if it is considered to be of higher quality.
The properties that are considered in RUMS can be reduced or extended as required. For example, it would be possible to compare the roles of the LEUs and add a partial indicator that indicates whether the two enterprise groups have the same GGH. However, after some empirical testing with the data from the use cases in chapter three, we found that it is more useful to analyse such a same GGH/different GGH similarity after calculating RUMS. It would also be possible to look more closely at the control relationships in the groups and see how many direct relationships are identical. This procedure would be much more complicated and in our point of view, the information on the group membership is more important than the individual relationship. The properties of employees and turnover are generally precise and complete in the registers considered, so that the intersections of the three properties LEUs, employees and turnover were considered in RUMS for the current use cases. This may be different in other registers so that these sub-indicators can either be deleted or replaced or the weighting can be changed by adjusting the parameters a-f.
Determining the weights is a very difficult question. We carried out simulation studies to analyse the effects of different weights. The populations were analysed to see how often several enterprise groups are similar and in how many of these cases there is no dominance of a comparison, so that the weighting is actually decisive for deciding which enterprise groups are more similar to each other. For these constellations, a grid was created for each parameter so that each parameter could assume 100 different values. 7 Each possible constellation was then simulated and a calculation was made to determine which comparison wins more often. In addition, a geometric procedure was used to determine the value of the parameters at which the decision changes. With both methods it could be shown that the equal weighting of the sub-indicators produces a result that is always a very probable result even with non-equal weightings. However, this weighting could be further investigated by applying these or other methods and also with more different use cases and data. We currently recommend an equal weighting of the parameters as long as there are no indications that one of the considered populations or properties is of poorer quality.
The use cases of RUMS can be separated in the necessity to compare groups from different points in time or to compare groups from different data sources.
Use case 1: comparison of groups at different points in time
At the beginning of every new reference year, the German NSBR receives data on enterprise groups. To process this data, a decision must be made for each existing enterprise group in the business register from the old reference year as to whether this enterprise group is to be continued and with which new group information of the new reference year. For this purpose, the enterprise group data of the reference years 2021 and 2022 were prepared and the RUMS-formula with equal weighting for the parameters a-f was applied. The enterprise groups at the two different points in time can now be divided into three parts (see Figure 3).

Comparison of enterprise groups between reference years 2022 and 2021.
Firstly, there are the identical enterprise groups in Figure 3 (the upper bars) with a RUMS = 1. For these groups, continuity is out of question since no changes concerning their group structure occurred in the two reference years.
Secondly, the lower bars of Figure 3 shows enterprise groups with no similarity to any of the enterprise groups in the other reference year, i.e. RUMS is 0. These are groups that are either completely new or no longer exist in 2022.
Lastly, the most interesting enterprise groups are in the middle bars in Figure 3. They show a varying degree of similarity with one or more enterprise groups from the other reference year. Their RUMS-values vary between 0 and 1. For these groups, RUMS can help to decide which new group information of reference year 2022 belongs to which group from the old reference year 2021.
With the help of RUMS, the German NSBR updated the structures of enterprise groups to the next reference year. For the enterprise groups with no similarity, the new enterprise groups were built up and the no longer existing enterprise groups were liquidated in the German NSBR. For the enterprise groups with a varying degree of similarity, RUMS can be used to identify the best match between the enterprise group of the current and the previous reference year. However, we also needed to determine how great a similarity between two enterprise groups must be in order to allow the assumption that this is the same enterprise group. In order to resolve this issue, we took an iterative approach. In the first step, only those enterprise groups with a RUMS > 0.9 were considered and continued. At each step, the limit for the RUMS-value was successively lowered and random checks were made to ensure that the results were still appropriate. If there were similarities of one group to more than one other group, the higher RUMS-value was used.
We carried out this iteration four times in total and ended at RUMS > 0.3. After performing some random manual checks with enterprise groups of low RUMS-values, we decided that the similarity between the enterprise groups of the different reference years is too small to continue the groups with the information from reference year 2022.
With the help of RUMS, the annual processing of new information could be automated to a large extent and many enterprise groups could be continued in a reasonable way, meaning that they keep the same ID.
Furthermore, RUMS can be an important indicator of either real changes between different reference years or differences in the quality of the data between different reference years. To find out the real causes, RUMS in combination with an indicator of the economic importance of the enterprise group can be used to select cases for the manual treatment of the most important enterprise groups.
For every reference year, all EU member states send their data on multinational enterprise groups of their NSBR to Eurostat. Here, the different data are combined in the EuroGroups Register EGR. 8 In a perfect world with perfect communication of the data, after the transmissions, every multinational enterprise group and all of the contained LEUs that are part of the German NSBR should be a part of the EGR. In practice, however, this is not the case. RUMS can be used to investigate to what degree the enterprise groups differ between both registers.
To investigate the similarity, the data of both registers were restricted to multinational enterprise groups with at least one LEU in Germany. In this use case only the German parts of the enterprise groups were analysed by calculating RUMS only with the German LEUs within the group.
We applied the RUMS-formula while using equal weights for the parameters a-f. In the end, we calculated a RUMS-value for all enterprise groups and the results can be divided into three parts again. Figure 4 shows the identical enterprise groups in the upper bars again with RUMS = 1. For these groups, the data transmission must have worked without any errors.

Comparison of enterprise groups between NSBR and EGR (reference year 2022).
Additionally, there are again enterprise groups with no similarity i.e. with RUMS = 0 (the lower bars in Figure 4). In the German NSBR, these are groups that do not exist at all in the EGR, meaning that their LEUs are not part of any groups in the EGR. In the EGR, these groups contain German LEUs that do not belong to any enterprise group in the German NSBR.
The last part again are the groups that are similar but not identical between both registers (the middle bars). The groups have RUMS-values between 0 and 1. These enterprise groups are similar to one or more enterprise groups of the other register.
RUMS and the division of the enterprise groups into the three parts could now be further used to perform quality checks for the most important groups. Concerning the groups with RUMS = 0, it could be analysed why specific groups from the German NSBR do not appear in the EGR (transmission errors, methodological differences). In addition, it would be interesting to take a deeper look at the groups in the EGR that are not part of the German NSBR at all. For these groups, it might be the case that the EGR has information about group relevant German LEUs that the German NSBR does not have.
Finally, with regard to the groups with 0 < RUMS < 1, for the economically most important groups it should be examined whether the group structure in the German NSBR is correct or not. So, in this case, the RUMS-value can also be used to select the groups that require manual quality checks.
The two use cases shown in chapter three demonstrate how RUMS quantifies the similarity between two enterprise groups and a deeper evaluation of all RUMS-values can further be used to analyse the similarity of all enterprise groups in two populations. Figures 3 and 4 give a visual indication of how similar the populations are. In other words, how similar either the data in two different registers are or how much changed over time within the same register. This can be quantified using RUMSA –
For this purpose, for both populations I and II, the maximum RUMS-values of all enterprise groups (most similar match with an enterprise group from the other population) are multiplied by the number of LEUs (L), employees (E) and turnover (T) in the respective group and then summed. These sums are then divided by the total sum of the properties L, E and T of the populations I and II. Population I contains n enterprise groups, population II contains m enterprise groups. The formula for this is:
Equation 4
In the first use case from chapter three, we have a RUMSA value of 0.951. The higher this value is, when comparing two points in time, the fewer changes can be found in the data. By calculating RUMSA between several reference years (2022 vs. 2021, 2021 vs. 2020 etc.) it would be possible to analyse in which reference year more or fewer changes have taken place in the structures of the enterprise groups, thus making the changes in the data measurable and comparable.
In the second use case, comparing data in NSBR and EGR, we have a RUMSA of 0.918. The higher this value, the more similar the data in the two registers are. Germany's aim is to increase this value over the next few years and to eliminate the reasons for the deviations in order to achieve a RUMSA value as close to 1 as possible.
There is an explicit global aim 9 to harmonise the statistical business registers of different countries and organisations worldwide in order to reflect the ongoing globalisation more consistently. For the first time, RUMSA enables the quantification of different information on enterprise groups in the registers. It also paves the way to measure possible improvements in the future.
Conclusions and outlook
As enterprise groups are often multinational and may have very large structures, it is increasingly important that their data are of high quality in the statistical business registers in order to be a good backbone for all business statistics. With the help of RUMS, enterprise groups as complex data objects are to be made comparable within statistical business registers, as long as they meet the requirements described in chapter two. This allows changes over time or deviations in different registers to be quantified and thus better analysed in order to increase quality. The RUMSA extension can be used to measure and compare the similarities of entire populations or registers. This enables a significant quality leap in the statistical business registers. Both similarity metrics can be adapted with regard to the properties of the enterprise groups in the populations and their weighting.
For further use cases, it will be interesting to examine the selected weighting of properties in more detail and to work out further optimisations or recommendations. In the current implementation of RUMS, it must be possible to match LEUs by identifiers. If this is not possible or if the overlap of the LEUs does not represent the similarity to be compared in other use cases, other metrics must be developed that could build on the idea of RUMS. For example, it might be interesting to look at enterprise groups from particular sectors in terms of specific properties and analyse them for similarities that are not based on overlapping LEUs.
The calculation of RUMS values has been programmed in the SAS software for the use cases presented in chapter three. This implementation could of course be made available. If there are more use cases from other organisations in the future, an implementation in R could also be supported to make this available in a standardised and more user-friendly way.
Footnotes
Acknowledgements
The authors would like to thank the editor and the reviewers at the SJIAOS. Special thanks go to Roland Sturm and Martin Beck for their valuable reviews and continuous motivation.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
