Abstract
Stats NZ’s Integrated Data Infrastructure (IDI) is a linked longitudinal database combining administrative and survey data. Initially, the IDI contained a small number of administrative datasets from key government agencies, which generally contained good quality identifying information such as names and date of birth. As a result, the methodologies developed to link these datasets together relied heavily on these variables, and yielded high link rates while maintaining good quality links. When survey datasets were later added, the link rates achieved were lower than that of administrative datasets, due to poor quality names and the underutilisation of geographic information. This indicated there were improvements to be made to the linking methodology used to link survey data in the IDI. Stats NZ underwent extensive consultation with the research community on their requirements for the expansion of the IDI (IDI2). A key finding from the consultation was that researchers wanted improved survey linkage. This paper outlines how the address history of individuals were used to increase the link rate of surveys in the IDI.
Introduction
This paper discusses how high-quality geographic information collected in surveys were utilised to improve the link rates of survey data in the Integrated Data Infrastructure (IDI). Section 1 introduces the IDI and the original methodologies and processes used to link survey data. It also presents the link rates achieved and the impetus to improve survey link rates in the IDI. Section 2 outlines the address-based linking methodology used for this improvement. We show the results of this method in Section 3, and conclude in Section 4 with a discussion and some learnings from this work.
Overview of the Integrated Data Infrastructure (IDI).
The Integrated Data Infrastructure (IDI) is a linked longitudinal database combining a variety of administrative and survey datasets [1]. Over the years it has become a primary database in New Zealand used for policy evaluation and academic research, and more recently for the production of statistical outputs, such as the 2018 Census [2] and child poverty statistics [3].
Data in the IDI is linked through a central dataset called the spine, which is comprised of the union of three core datasets – births registered in New Zealand from the Department of Internal Affairs (DIA), tax registrations from Inland Revenue (IR), and visa applications from the Ministry of Business, Innovation, and Employment (MBIE) [4]. The spine aims to represent an ever-resident population for New Zealand (the population of individuals who have ever lived in New Zealand), although under and over coverage can occur [1, 4]. All other datasets are linked to the spine. Links across datasets are therefore deduced where records from different datasets are linked to the same spine record (see Fig. 1).
The linkage of datasets in the IDI is separated into two phases. Phase 1 involves linking administrative datasets to the spine, such as health, social services, education, and justice data. After phase 1 linking is finished, the spine is ‘rebuilt’ to include geographic variables from the datasets linked in phase 1. Phase 2 linking then uses this geographic information to link the survey datasets. As address is the mechanism in which we identify and contact survey respondents, geographic information is high quality and is non-missing in surveys.
Fellegi and Sunter’s probabilistic linking [5] is the main method used to link data in the IDI, whereby weights that represent the probability that two records belong to the same person are assigned to record pairs [6].
Sometimes, deterministic linking or exact matching is used in combination with probabilistic methods. This occurs when common unique identifiers are available (e.g. tax numbers from IR), or where exact matching on demographic variables (e.g. name, birthday, sex) is incorporated. This applies to both phase 1 and phase 2 linking.
Link rates and false positive rates achieved for 2013 Census and survey datasets in the IDI, based on September 2018 data
Link rates and false positive rates achieved for 2013 Census and survey datasets in the IDI, based on September 2018 data
Two phases of linking in the Integrated Data Infrastructure (IDI).
An example based on synthetic data of how address history is included in the spine.
Currently, the datasets linked during phase 2 are the 2013 New Zealand Census, Household Economic Survey (HES), Household Labour Force Survey (HLFS), General Social Survey (GSS) and the Programme for the International Assessment of Adult Competencies (PIAAC).
Initially, the IDI only contained data from administrative sources and thus only needed one phase of linking. These datasets generally contained good quality identifying information such as names and date of birth. The methodologies developed to link these datasets together relied heavily on these variables, and yielded high link rates while maintaining good quality links.
The two-phase approach was developed when 2013 Census data was added to the IDI. To optimise the linkage of 2013 Census to the IDI spine, we needed to incorporate geographic information into the linking methodology. The three datasets that comprise the spine lack high quality and high coverage geographic information on their own, so the two-phase process was established to attach geographies contained in other administrative data that had already been linked to the spine (see Fig. 2).
As one spine record can be associated with multiple geographies in the phase 1 administrative datasets across time, the meshblock closest to Census night 2013 was extracted and added to the spine. This ensured the time dimension of the spine geographies was consistent with the geographies collected in the 2013 Census. The achieved link rate was 94%, with a false positive rate of 1.5%.
When other survey datasets were later added, such as HES and HLFS, the linking methodologies applied also utilised geographic information to optimise the link rate and the quality of the links produced. This was necessary as the quality and availability of names collected in these surveys were not as high as in administrative datasets. However, as the geographic variables on the spine were time referenced to 2013, the link rates achieved were not as high (Table 1), as the addresses in these surveys span many years, and are not necessarily close to 2013.
These relatively low link rates indicated the potential to increase the link rate by revisiting the linking methodology for survey data. Stats NZ underwent extensive consultation with the IDI research community on their requirements for the expansion of the IDI (IDI2). A key finding from the consultation was that researchers wanted improved survey data linkage in the IDI.
Additionally, as New Zealand continues to move towards the combined use of survey and administrative data to produce Official Statistics and to provide insights, the high quality linkage of these datasets is of utmost importance. Combined use of survey and administrative data helps remedy known difficulties in conducting surveys, such as declining response rates, missing information, measurement error, and respondent burden [7].
Recently, as a response to the lower than expected response rate for Census 2018, Stats NZ used administrative data from the IDI to add non respondents to the census dataset [2]. Additionally, Stats NZ’s measurement of child poverty used income from administrative data to replace income collected from survey respondents, as survey reported income is often a rough estimate or not disclosed [3].
The combined use of survey and administrative data has also been explored in other countries. The English Longitudinal Study of Ageing (ELSA) links to administrative data on insurance, benefits, tax, mortality and cancer in order to better understand factors that influence the ageing process [7, 8].
Additionally, national record linkage centres have been established in the UK and Germany [9, 10], a whole-of-government data integration initiative in Australia [11], and the Italian National Statistical Institute has built a data integration tool in house, with support from their Spanish counterpart [12].
Address is used as the primary mechanism for identifying and contacting survey respondents at their household. Primary sampling units are geographic, where dwellings in selected units are enumerated by field staff. The interviews take place at the address. For these reasons, address information is high quality and non-missing in surveys.
In comparison, we ask survey respondents to provide their name, but it is not compulsory to do so. Respondents may provide partial names, nicknames, fake names, or none at all.
Therefore, the strategy undertaken to improve survey link rates in the IDI was to move to an address history-based methodology. We maximise the use of address in the linking methodology to allow less reliance on strict name agreements.
Improvements to the spine
To enable this, the time dimension of geographic information in the spine will no longer be fixed to a specific reference date, i.e. that of 2013 Census night. We include the entire address history of individuals in the spine. That is, every unique address associated with a spine record post phase 1 linking is included when rebuilding the spine for phase 2 linking.
Original linking methodology used to link HES and HLFS to the IDI spine
Original linking methodology used to link HES and HLFS to the IDI spine
Address based linking methodology used for linking survey datasets to the IDI spine, using full address history
This means that if there are three unique address notification records for an individual in the administrative datasets, the record for the individual in the spine is repeated three times, one for each different address. The addresses are also time referenced to the first and last date in which there is an address notification in the administrative data.
In the example shown in Fig. 3, a spine record for Bob Smith has geographic information in the health and education datasets, which are identified after phase 1 linking. The old process was to attach the geography time referenced closest to 2013 census night to the spine. In this example, this refers to geographic ID 350078.
The new process attaches all the unique geographies for Bob Smith’s record to the spine, so that his record is repeated in the spine for each different geography with the relevant time reference. First notification date is the earliest date we have a record for Bob at the address, and last notification date is the latest date.
Additionally, the smallest geography in the spine was meshblock. The most precise geographic variable available in the IDI is the Address ID. This is the raw address string transformed into an ID. The assignment of these ID’s occurs during address geocoding, in which all IDI addresses are linked to Stats NZ’s Statistical Location Register (SLR). Multiple IDI addresses that are linked to the same SLR address are assigned the same ID, therefore creating a common address ID across all datasets. We now also include the Address ID, as this variable can be used to increase the strength of the links created.
Survey data was linked using geographic variables such as meshblock. However, this method was limited because only the meshblock associated with each person that best represented their address on 2013 census night was included on the spine. On the other hand, survey addresses are time referenced to the date of collection.
Non-geographic passes were used to account for this, but these rely heavily on names and date of birth (see Table 2). However, respondents often provide nicknames, fake names, or do not agree to disclose identifying information as they prefer to retain a degree of anonymity.
Having access to geographic histories for individuals in the spine gives us the full range of addresses across time to create links between spine and survey records.
Table 3 outlines the new address-based linking methodology. Time dimension of address is captured by the reference period. As the address timestamps in the administrative data are subject to time lags, they are not absolutely indicative of when the individual resides at the address. We transform the first and last address notification year into a time interval that represents when they could have lived at the address. If the year of the survey interview is contained between that interval, then there is agreement in the reference period of address variable. This agreement will increase the overall weight assigned to the record pair, therefore increasing the probability the record pair refers to the same person.
Pass 1 filters possible links to records with the same date of birth and the same address, at any point in time. If the records also have either similar names, or agreement in the reference period of address, then the link is accepted. Name is included in this pass to accept links where record pairs have agreement in address, names, and date of birth, but disagreement in the reference period of address.
Pass 2 repeats pass 1, however does not include names as a linking variable. It accepts links where two records have the same address and date of birth, and there is agreement in the address time reference. It does not matter if there is total disagreement in first and/or last name (wherein the first pass this may have resulted in an unaccepted link). This is designed to accounts for survey responses with missing or fake names.
Improvements in link rate for HES and HLFS, September 2018 vs March 2019 – where September 2019 uses the old methodology, and March 2019 uses the address based methodology
Improvements in link rate for HES and HLFS, September 2018 vs March 2019 – where September 2019 uses the old methodology, and March 2019 uses the address based methodology
Percentage of total links by pass for HES and HLFS under the address history based methodology (March 2019)
Improvements in HES link rate by year of birth and sex, September 2018 vs March 2019.
Pass 3 filters possible links to be records who have ever lived at the same address. The linking variables include names, date of birth, sex, and the address reference period. A bespoke cut off must be chosen here to only accept links with enough agreement in the linking variables.
Additional non-geographic passes are added to account for records with no address information in the administrative data, or where there are linkage errors from the geocoding process that prevents address matches being made across datasets. For example, pass 2 and pass 4 outlined in Table 2.
This address based methodology was implemented in the IDI as part of the March 2019 production rebuild. It was applied to HES, HLFS, GSS, and PIAAC. For simplicity, we present the results for HES and HLFS from March 2019, where the address based methodology is used, and compare it to the September 2018 production results, where the old methodology was used. GSS and PIAAC show similar results.
Adding address histories to the spine increased the number of rows from 10 million to 38 million. We therefore hold an average of 3.8 different addresses for an individual in the phase 1 administrative data.
As shown in Table 4, the link rate for HES increases from 78.2% in September 2018 to 94.8% in March 2019, and for HLFS increases from 81.5% in September 2018 to 91.6% in March 2019. Additionally, the quality of the links produced remain similar.
Table 5 shows that more than 80% of the links for both HES and HLFS are made in the first pass. This is where address and date of birth are an exact match, and the only requirement is that the weight assigned for the record pair is positive, meaning that either the reference period of address overlaps, and/or the names are similar (and not necessarily both). Links picked up in pass 2 tend to be those with completely different names (which may represent false positive links), respondents with missing names, respondents that use family roles for names such as “father” or “son”, and records where the first and last name are transposed. For HES, a very small number of links were made in his pass, so it may be removed in the future.
Improvements in HLFS link rate by year of birth and sex September 2018 vs March 2019.
Improvements in link rate by year of collection for HES, September 2018 vs March 2019.
Improvements in link rate by quarter of collection for HLFS, September 2018 vs March 2019.
The address based linking methodology improves link rates for females and males across all ages. Figures 4 and 5 show this improvement, as well as the stabilisation across year of birth, and a closed gap between the female and male link rate. There is a noticeable dip for the HES and HLFS link rate for both males and females who are born between 1985 and 1995, and this dip seems to remain under the new methodology (Figs 4 and 5). This could be attributed to the difficulty in the capture of address information for young adults in our administrative data. They tend to be more mobile and are more likely to have administrative addresses that are out of date [13].
Figure 6 shows dips in the September 2018 HES link rate for the earliest and latest collection years, corresponding to where the addresses collected are furthest away from Census night 2013. For the years 2016 and 2017, HES stopped collecting date of birth for respondents under 15 years of age, resulting in an additional decrease in link rate. Figure 7 shows the September 2018 HLFS link rate is the highest around 2013, again a by-product of the spine geographies being time referenced to 2013 Census night.
The new address history based methodology increases the HES and HLFS link rate for all years and quarters, and is more stable across time as the entire address histories are included in the spine. Using address reduces the effects of decreased quality and increased missingness in name information, as it allows links to be created regardless of strict name agreement.
The results from the March 2019 IDI production run have shown the address history based method significantly improves survey link rates in the IDI, providing researchers with a more complete picture of the survey population.
Going forward, we will continue to refine the address history based method. This includes how to best utilise the combination of address, meshblock, and other higher-level geographies in the linking methodology. Additionally, any new datasets added to the IDI that contains address information will be able to be linked using this method.
As the spine is the central dataset in which all other datasets are linked to in the IDI, it is of utmost importance that it is built in a way that is flexible and future proof to datasets that may be linked now and in the future. We must ensure that variables contained in the spine are not restricted in any way to favour linkage of one dataset and not another.
Many Official Statistics agencies are moving towards automated production processes, and it may be desirable to apply standard linking methodologies to link data in the IDI to reduce human intervention. However, the results in this paper demonstrate the importance of designing linking methodologies tailored to the quality of variables in the dataset being linked, and this requires the expertise of data linkers and data suppliers.
As Stats NZ increases the use of administrative data in producing Official Statistics, there is a heavy reliance on high quality linked survey and administrative data. Additionally, surveys will continue to play a vital role in the IDI, providing information that simply cannot be captured in administrative data– such as measures of wellbeing. The successful integration of survey and administrative data is therefore critical in not only the future of Official Statistics, but also in our ability to understand social and economic issues in New Zealand.
Discussion
The issues discussed in this paper highlights the dependence on high quality linking variables for successful data integration to be achieved. High missingness or low quality in core variables such as name, date of birth, or address significantly effects the linkage results. In cases where datasets with poor linking variable quality and availability have been linked in the IDI, link rates and link quality have also been poor. If the source agency has a vested interest in the research to be done on the linked data, the poor linkage results can be a precursor to improved data collection. In relation to this, the Stats NZ Chief Executive was recently appointed as the Government Chief Data Steward [14]. As part of this role, Stats NZ are leading the establishment of Data Content Standards to encourage consistency across government in the content and quality of key linking variables [15].
The Data Content Standards work enables an alternative or complementary approach to improving survey link rates in the IDI, where we mandate improved collection of core linking variables. Additionally, the collection of agency unique identifiers that assist linking (for example IR tax number) could be incorporated. An experimental UK study asked respondents in the Improving Survey Measurement of Income and Employment survey whether they consented to the linkage of their data to administrative records, and if so, to provide their National Insurance Number (NINO) to assist this. About 78% of the sample provided consent, and 89% of the consenting respondents provided their NINO [7].
However, the primary objective for surveys is the collection of accurate data on the subject matter of interest (for example wellbeing, expenditure, or income). Heavy emphasis placed on providing accurate name information could compromise respondent trust in terms of confidentiality and anonymity, so trade-offs should be discussed.
The work discussed in this paper was achieved through the collaboration of data linkers, subject matter experts, researchers, and applications developers to address a pressing issue faced by the IDI research community. This researcher-producer feedback loop should be used going forward, as Stats NZ continues to develop and improve the IDI in order to unleash the power of data to change lives.
Glossary
Link rate – the number of unique records in a file that are linked to the spine, divided by the total number of unique records in the file.
False positive rate – two records that were linked in error and do not correspond to the same unit [16]. False positives in this paper are estimated through clerical review or a predictive logistic model.
Pass – one iteration of a record linkage process, using a particular set of blocking and linking variables [16].
Linking variable – variables used to compare two records to see how likely it is that both belong to the same unit [16].
Blocking variable – variables that are used to divide a file in a block. Blocks are groups of files to be linked that have some information in common. Records are only compared with others in the same block. Using blocks reduces the number of comparisons that must be made [16].
Cut off – the weight at or above which all record pairs are linked and below which all record pairs are not linked [16].
Soundex – a phonetic encoding algorithm that writes a string of characters based on the way the string is pronounced. For example, ‘Camden’ and ‘Comden’ both encode to C535 [16].
Meshblock – the smallest geographic unit for which statistical data is reported by Stats NZ [17].
Footnotes
Acknowledgments
I would like to express my very great appreciation to my fellow colleagues at Stats NZ who kindly provided me comments that helped improved this paper. The implementation of this work would have not been possible without the support and expertise of those in the Integrated Data production team and the applications developers, both of whom run the engine room of the IDI.
