Abstract
The cohort-component method is the standard model for producing population projections in official statistics. It is straightforward to compute, requires minimal input data, and is widely recognised by demographers. However, cohort-component projection models are limited in their ability to capture complex demographic processes and provide detailed output for individual-level outcomes. To address these shortcomings, Statistics Austria has developed the dynamic microsimulation model STATSIM for its official national population projection, which was previously computed by the cohort-component method. We have opted for a gradual transition, starting by replicating the results of past cohort-component projections using microsimulation. As a first extension, we have implemented a model of international migration that takes into account the relationship between emigration risk and length of stay, as well as country of birth. By comparing the results of STATSIM's retrospective projections with counterfactual cohort-component projections, we show that STATSIM's projections are more consistent with observed emigration patterns. In the future, STATSIM can be further developed by adding modules for education, employment, health and other socio-economic characteristics.
Introduction
Before 2022, Statistics Austria produced its national population projection according to the cohort-component method, a deterministic method that projects the components of population change separately for each birth cohort using event rates. Indeed, the cohort-component method is the standard tool for the production of population projections in official statistics, governmental institutions and international organisations. 1 It is well established in the literature, closely linked to demographic theory and mathematically simple. However, the cohort-component method cannot model complex demographic processes, account for multi-faceted population heterogeneity or produce results for a variety of individual-level characteristics. 2 Microsimulation presents a solution to these issues. As formulated by Orcutt in 1957, microsimulation studies in the social sciences can be seen as a counterpart to experiments performed in the natural sciences. 3 The present article focuses on dynamic microsimulation. In this class of models, individual life-courses are simulated over time: 4 Actors can move from one province to another, start school or employment, give birth to children, and so on, until they either die, move abroad or reach the end of the projection horizon. The simulated individuals and events are then aggregated, so that the development of the macro-system – in our case, a population – can be analysed and projected into the future. Hence, while microsimulation models behaviour at the individual level, it can be used to describe a macro system.2,4 The underlying idea is that in the presence of heterogeneous agents, individual behaviour is easier to understand and model than aggregate behaviour. 5 Unlike the cohort-component method, microsimulation can handle large sets of individual-level characteristics efficiently. It is also better suited to dealing with interdependent demographic processes, as it provides the necessary flexibility to model interactions between characteristics (e.g. education and fertility) as well as individuals (e.g. mother and child). 2
Following Orcutt's seminal paper, 3 the first dynamic microsimulation was presented in Orcutt et al. 6 Given the inclusion of core demographic processes, such as births, deaths and migration, in most dynamic microsimulation models in the social sciences, there is a wealth of literature available for those seeking to produce population projections. For an overview of different applications of (dynamic) microsimulation as well as general developments and challenges in the field, see Li and O’Donoghue, 7 O’Donoghue, 8 O’Donoghue and Dekkers 9 and Zaidi et al.; 10 for a more recent article on applications in Germany and Great Britain, see Schnell and Handke. 11 Within official statistics, several National Statistical Institutes (NSIs) use microsimulation. Perhaps the most prominent is Statistics Canada, which has developed a range of models, including the population health model POHEM 12 and Demosim, 13 a microsimulation model for detailed population projections. In Europe, Statistics Norway employs MOSART for analyses and projections regarding education, labour supply and public pension spending 14 and LOTTE for tax-benefit simulations. 15 INSEE, the NSI of France, performs pension policy analyses with its dynamic microsimulation model Destinie.16,17 More recently, the research group MikroSim, which has affiliations with the Federal Statistical Office of Germany, has developed a regional dynamic microsimulation model for Germany. 18 For a survey of microsimulation models used by public bodies in the EU, including NSIs, see Dekkers and Van den Bosch. 19
Comprising more than one in five Austrian residents (22.3% as of January 1st 2024), first-generation immigrants constitute a highly heterogeneous group, including, for example, university students from Western Europe, workers from East and South-East Europe and refugees from the Middle East. 20 However, most population projections treat migrant populations as fairly homogenous. 21 A microsimulation-based population projection model that considers the heterogeneity among immigrants and produces results for Austria and the other EU member states as well as Switzerland, Norway, Iceland and the United Kingdom is QuantMig-Mic. 22 The model contains a range of migrant-specific characteristics, including length of stay, place of birth and age at immigration. Furthermore, interactions of these variables with education and labour market outcomes as well as fertility are considered. Other applications of dynamic microsimulation that include region of birth and length of stay as immigrant-specific characteristics are presented in Marois et al. 23 and Potančoková et al.. 24
In order to model the mechanics of population change in Austria in a more realistic way and to produce richer projection output, Statistics Austria has developed the dynamic microsimulation model STATSIM. Since 2022, STATSIM is used to produce the national population projection of Statistics Austria. Moving from the cohort-component method to microsimulation represents a fundamental methodological change. To ease the transition process and enable users to track changes step-by-step, we began by replicating the results of Statistics Austria's 2013 population projection, which was computed using the cohort-component method, through microsimulation. As a first step in extending STATSIM beyond the cohort-component framework, we incorporated a more detailed model of international migration. It considers immigrants’ country of birth and duration of stay as determinants of emigration behaviour, as emigration risks differ substantially across sub-groups of the migrant population and decrease at the individual level with the length of stay. 25 We chose the topic of international migration for the first model extension because the future size and composition of the Austrian population are significantly influenced by the demographic behaviour of its foreign-born population – a situation that is becoming increasingly common among European countries with low fertility and high levels of immigration. 25 In the future, STATSIM can be developed further, with additional modules for education, employment, health and other socioeconomic characteristics.
The aim of this paper is twofold. Firstly, it discusses why and how Statistics Austria transitioned its national population projection from the cohort-component method to microsimulation. Secondly, it presents the STATSIM model and uses retrospective projection to compare it to Statistics Austria's previous projection model. Section 2 provides a theoretical comparison of the cohort-component method and microsimulation for population projections, gives a brief overview of the corresponding literature and discusses the issue of uncertainty. The features, dynamics and structure of STATSIM are presented in Section 3. Input data is covered in Section 4. Section 5 contrasts STATSIM and the cohort-component method by comparing retrospective projection results and assesses the impact of the methodological change on projection accuracy. Section 6 concludes.
The cohort-component method vs. microsimulation for population projections
The standard cohort-component method derives from the population balancing equation: Population change from time t to t + 1 is computed by taking the population level at time t, adding births and net migration and subtracting deaths, which occurred between t and t + 1.
26
Microsimulation offers a solution to these limitations: Unlike macro-level projection models, such as the cohort-component method, microsimulation builds on the characteristics of individuals instead of cohorts. Data storage and updating therefore pertain to vectors of characteristics for each simulated individual instead of the full cross-classification table. As long as some degree of independence between individual-level attributes can be assumed, microsimulation models can process large sets of characteristics more efficiently, allowing for greater model complexity as well as flexibility. 2 , i Another advantage of microsimulation is its ability to account for Non-Markov processes: 4 As the model keeps track of individual life paths, past states can affect future states. In cohort-component models, the information about how a person got into a cell of the cross-classification table is lost. It is therefore not possible to model, for example, that individuals who have experienced frequent relocations or have resided in a particular region for a short time exhibit higher mobility. For a detailed discussion of the strengths and weaknesses of microsimulation for population projections, especially when compared to the cohort-component method, see Van Imhoff and Post. 2
While the cohort-component method has remained the standard method for producing population projections, advances in computing power, an expansion in the availability of micro data and a growing literature on bridging the gap to microsimulation are making the transition more appealing.4,26,30 For example, Jia et al. replicate Statistics Norway's multi-regional cohort-component projections by microsimulation. 30 They demonstrate that for aggregate values, such as the regional population level, microsimulation results converge to those of cohort-component models after a small number of model runs. After a thousand simulations, the results of the two methods are virtually identical, even for subgroups of the population. 30 In another recent article, Puga-Gonzalez et al. try to reproduce the United Nations’ cohort-component projections for Norway using microsimulation. They document the challenges faced in the process and highlight the implicit assumptions of the cohort-component method, which become explicit in microsimulation modelling. 26 Following on, Bacon et al. discuss model design features that are necessary to recreate UN cohort-component projections for Norway, the United States and India using microsimulation. 31 Taken together, the literature shows that microsimulation can produce the same results as cohort-component models but it can also move beyond, providing greater flexibility as well as transparency in modelling demographic processes.
As described by Imhoff and Post, “[a]ny projection into the future is subject to random variation”. 2 This holds true for deterministic macro models, such as cohort-component models, as well as microsimulation. It is therefore worth discussing the issue of uncertainty in these two model types. For a given a set of parameters, cohort-component models compute the expected value of the population, while microsimulation produces realisations of a random variable which, on average, converge on the expected value as the number of model runs and/or the size of the simulated population increase. 2 Each individual simulation run is subject to Monte-Carlo variability, i.e. variation introduced by drawing random numbers. Jia et al. refer to this as ‘model-specific uncertainty’, 30 Imhoff and Post call it ‘inherent randomness’. 2 This type of uncertainty is larger for small subgroups of the population. 30 By repeating the simulation multiple times, we obtain an average result as well as a range of realisations, from which we can compute standard errors. This measure of uncertainty is most useful when simulating the actual population size: If assumed probabilities are accurate, it expresses the expected randomness. To reduce Monte-Carlo variability to a negligible level in Statistics Austria's national population projection, we run STATSIM with a sample size of ten times the actual population.
Following Kennedy and O'Hagan's classification, other types of uncertainty in computational models include uncertainty in parameters; model inadequacy, i.e. the difference between the mean outcome of the real process and the model prediction, even when parameter values are true; measurement error and other forms of uncertainty in observations; parametric variability, which is introduced by drawing parameter values from a distribution rather than using set values; and residual variability, which can result, for example, if the true process is stochastic and hence uncertainty remains even when the model and its inputs are correctly specified.32,33 In population projections, a particularly important aspect is parameter uncertainty, i.e. uncertainty regarding the underlying assumptions about future demographic behaviour. In classic deterministic cohort-component projections, a range of variants is usually produced to illustrate this type of uncertainty. 26 This approach has been critiqued, as deterministic variants contain no quantified information about their likelihood. 34 Probabilistic cohort-component projections have been presented as a solution. 35 These models quantify parameter uncertainty by taking random draws from the probability distributions of the parameters and producing cohort-component projections from each set. Combining the results of these projections provides a probability distribution over the outcomes of interest, such as population size. 35 Some NSIs use this method to measure variability in their official projections (e.g. the Netherlands 36 and New Zealand 35 ). In principle, the same approach can be used in microsimulation-based projections, by making transition rates random. This is not currently done in STATSIM, but would provide an interesting field for further development in the future.
Furthermore, it is useful to note the issue of ‘specification randomness’: As discussed by Imhoff and Post, each simulated event introduces another random draw and each added characteristic requires estimation of a greater number of parameters, which are themselves subject to randomness. 2 Hence, the inherent randomness increases with model complexity, resulting in a trade-off between low variability and detailed modelling of demographic processes. On the other hand, failure to account for relevant relationships between variables introduces model specification error, which leads to biased projection results. In Kennedy and O’Hagan's terms, it would be a case of ‘model inadequacy’. 32 Hence, microsimulation models should have a sufficient level of detail to be correctly specified, without becoming so complex that projection accuracy becomes an issue. At present, STATSIM is a relatively simple model with a small number of characteristics, so specification randomness is not a problem. Not capturing important dynamics of demographic behaviour, such as the duration-dependency of migration, would present a far bigger issue in the form of specification bias. For a further discussion of uncertainty in population projections, see Van Imhoff and Post, 2 Jia et al. 30 and Lutz and Goldstein34.
Methodology
Model features
According to Spielauer, 37 methodological aspects of microsimulation models can be grouped into three dimensions: (1) the way a simulated population is created, (2) the modelling of time, (3) the simulation schedule, i.e. the order in which individuals are simulated. Regarding the simulated population, we can distinguish cross-sectional and synthetic starting populations as well as open and closed population models. Time can be treated as a discrete or continuous variable. Lastly, in terms of the simulation schedule, we can differentiate non-interacting and interacting population models. In the following, we give a brief overview of these features, focusing on the approaches used for STATSIM. For a more detailed discussion, see Spielauer. 37
Population
The population at the beginning of the simulation is referred to as the ‘starting’ or ‘base’ population. In general, there are different approaches to creating a starting population, the choice usually being based on available data sets (survey vs. administrative data), the need to impute characteristics from different data sources, and confidentiality issues. In dynamic microsimulation, cross-sectional starting populations are the most common type. The starting population in STATSIM is cross-sectional and produced from individual-level register data. In contrast, a simulation can also be based on a fully synthetic population, in which the life histories of all individuals are simulated from the moment of birth, so that all past events are modelled. An example of this approach was the Canadian LifePaths model. 38
STATSIM is an open population model in the sense that immigrants are added to the simulated resident population and residents can leave the simulation by emigration. There is no partner matching in STATSIM. In models with partner matching, the distinction between open and closed population models has another meaning: In closed population models, partners have to be searched for in the existing population. In open population models, they are created on demand and treated as ‘attributes’ of their spouse. 37 For more information on partner matching, see Zinn. 39
Time
Microsimulation models can treat time as a discrete or continuous variable. In continuous time models, events can be realised at any point in time and are simply aggregated by simulation year to produce results for each projection year.7,37 In discrete time models, states can only be updated once in a given, fixed time interval (e.g. once a year). The precise moment within the interval at which underlying events occur as well as information on potential multiple events, such as the number of events, their sequence and the time span between them cannot be determined. In continuous time models, these limitations are overcome through the use of competing risks, 40 which are discussed in Section 3.2. STATSIM uses a continuous time framework, allowing us to explicitly model the seasonality of migration (see Section 3.3.4) and produce outputs at any reporting date within a calendar year. Continuous time adds considerable flexibility to the model, as many socio-demographic processes have seasonal dimensions or specific schedules, such as school years and semesters. Potential drawbacks of continuous time models include greater complexity and computational intensity, due to event queuing, increased data demands and availability of appropriate software. In STATSIM, these limitations are mitigated through the use of administrative data sources, which provide daily information on demographic events, as well as the microsimulation programming language Modgen, which was developed at Statistics Canada and handles continuous-time simulation efficiently. 41 For a further discussion of continuous time microsimulation and its advantages compared to discrete time models, see Willekens. 40
Simulation schedule
A further classification of microsimulation models is whether they simulate one case at a time from start to finish (‘non-interacting’ or ‘case-based’ models), or the whole population simultaneously over a given time interval (‘interacting’ or ‘time-based’ models). Here, a ‘case’ refers to a simulated individual and any children, grand-children etc. generated in the course of the simulation. In case-based models, interactions can only take place between individuals of the same case. While this is a strong limiting factor for the kinds of events and processes that can be modelled, it constitutes a major computational advantage, both concerning speed (e.g. execution can be parallelised) and memory requirements (it is not necessary to have the whole population present, cases can be processed one by one). Time-based models are typically produced with population samples of less than one million persons due to the large computational effort of running this type of simulation. 37 In STATSIM, we want to utilise the large, administrative datasets available at Statistics Austria, which cover the entire Austrian resident population (9.2 million as of January 1st 2024 42 ). Hence, STATSIM is implemented as a case-based model.
Model dynamics
In terms of model dynamics, STATSIM follows a competing risk approach: At the start of the simulation, each person is assigned randomly drawn waiting times, given their individual characteristics, for all possible events. The event with the shortest waiting time is realised, once the waiting time has elapsed. This process is based on the idea of ‘stochastic race’, which posits that a more advantageous option has a higher likelihood of being selected in a given time frame, resulting in a shorter average waiting time and a higher probability of selection. 43 As soon as an event occurs, the characteristics of the simulated person are updated. Based on the new characteristics, new waiting times are assigned for all events. If a person dies, moves abroad or reaches the end of the projection horizon, the simulation ends for this person and the next person is simulated. For a detailed discussion of competing risks, see Galler. 44
Figure 1 demonstrates these dynamics for a model with three events: internal migration, death and the person's birthday. The lengths of the three horizontal lines at the top indicate the waiting times for each event at the beginning of the simulation, given the person's characteristics at that point in time. Initially, the waiting time until the person migrates to another province is the shortest, meaning that this event will occur once the waiting time has elapsed. When the event occurs, the province of residence of the person changes and waiting times are updated accordingly. In this example, the waiting time until death changes, as it is contingent on the province of residence. The waiting time until the person's birthday does not change, as it is a fixed date that is not contingent on other characteristics. The new waiting times are demonstrated by the three horizontal lines in the lower part of the figure. Now, the birthday event has the shortest waiting time. Once the waiting time has elapsed, the event occurs and the person's age is increased by one year. All waiting times will again be updated given the change in age, and so the simulation continues.

Model dynamics of a competing-risk dynamic microsimulation model with three events.
In dynamic microsimulation, individual life-courses are simulated over time. Since our aim is to produce a population projection, we focus on simulating core demographic events: births, deaths and migration. For each component, there are individual modules, which contain the corresponding event functions. The content of these modules is presented in sections 3.3.1 to 3.3.7.
The core characteristics in the model are age (0, …, 100+), sex (male/female), province of residence (nine federal provinces of Austria) and dichotomous country of birth (native/foreign). The international migration module uses duration of residence as an additional characteristic and replaces the dichotomous country of birth with detailed country clusters.
Mortality
In cohort-component projections, assumptions about future trends in fertility, mortality and migration are usually specified in terms of demographic rates, i.e. the number of events that take place over a set period of time relative to time spent at risk by the population of interest.
45
In the case of mortality, for example, the number of male deaths in a given year t,
In the current version of STATSIM, the fertility hazard
When a birth event occurs, the child is created as a new actor in the simulation, its sex is randomly drawn given the empirical distribution and the province of residence is passed on from the mother. The latter constitutes the only type of between-individual interaction currently implemented in STATSIM. In the future, this could be extended, for example, to passing information on the mother's country of birth or her level of education to the child.
Internal migration
For internal migration, we first determine the waiting time until a person emigrates from their current province of residence. Waiting times are obtained on the basis of migration probabilities, defined by age, sex, province of residence and dichotomous country of birth. If the event occurs, a destination province is sampled, given the individual characteristics. At present, the modelling of internal migration still follows the cohort-component logic in the sense that each individual can only move between provinces once a year. In the future, this can be extended by using hazard rates that depend on the duration of residence in the current province and allow for multiple internal migration events within a projection year.
International migration
The first modules implemented in STATSIM that add detail relative to the previously used cohort-component projection model relate to international migration. Among the 9.2 million people living in Austria at the start of 2024, more than 2 million were born abroad.20,42 When comparing groups of immigrants from different origin countries, there are variations in age and sex structures, employment status, education level, family composition, and duration of stay, among other things. Using data from the Austrian Migration Statistics on persons who immigrated between 2017 and 2019 – three relatively stable and representative years in terms of migration to Austria – the share of females among the 10 countries with the highest numbers of immigrants ranges from 36% for migrants born in Poland to 50% for native Slovaks. The share of young people aged 20 to 24 varies between 13% among native Bulgarians and Slovaks and 23% among Italian migrants. Correspondingly, data from the Register-based Labour Market Statistics show that the share of migrants enrolled in tertiary education within the first few months following immigration ranges from 1% among Romanian migrants to 23% among native Italians. To account for this heterogeneity in the population projection, we developed a model of international migration that incorporates detailed information on the country of birth.
Accounting for country of birth
Country of birth is an important factor in the emigration behaviour of migrants as it relates to the initial reasons for immigration and the costs of (return) migration.46,47 As immigration numbers from many countries are too small to model them individually, we establish clusters of countries whose immigrants show similarities in their emigration behaviour. The statistical clustering exercise is presented in the Appendix.
The results are intuitive: We identified clusters of predominantly work-related migration (Eastern EU) and mainly education-related migration (mostly Western EU). We also established three clusters of countries with large proportions of asylum-seeking migrants (mostly from the Middle East and Africa), distinguished by length of stay. Among the three, Syria constitutes its own cluster, as the duration of stay of Syrians in Austria by far exceeds that of migrants from other countries with large numbers of asylum applications. Figure 2 shows the country clusters in Europe, with each pattern corresponding to a different cluster.

Country clusters in Europe. Notes: Each pattern/colour corresponds to a different cluster.
For native Austrians, the federal province in which they were born is used in place of the country clusters.
Given the results of the cluster analysis, we formulate assumptions for future immigration in absolute numbers by country group for each projection year. In the microsimulation, immigrants’ age, sex and province of residence at the time of immigration are sampled from empirical distributions for each cluster. The month in which immigration occurs is also sampled, given the individual's broad age group ii and country cluster, to capture the seasonality of migration. In our continuous time framework, this enables us to derive more appropriate population estimates for different time points within a year, such as the beginning of a quarter or a school year/semester. Parameters are computed from individual-level migration records, available at Statistics Austria, which contain the exact dates of migration events for anyone who has been registered as a resident in Austria for at least 90 days. These records are derived from the central population register (in German: ‘Zentrales Melderegister’). For the 2022 round of population projections, data were pooled over the years 2017 to 2019 to derive stable empirical distributions for the input parameters.
International emigration
Unlike births, deaths and internal migration, international emigration is not determined on the basis of simple demographic probabilities or rates. Instead, we want to account for the relationship between emigration risk and time-dependent covariates, particularly length of stay. To this end, we estimate piecewise constant emigration hazards for each sex and country cluster. We first perform episode-splitting on the migration data, to divide the observation period for each individual into intervals with constant covariate values, creating a new interval whenever an event occurs that changes the value of one of the covariates. Transforming the data in this way allows us to examine how the risk of emigration changes over the life course, e.g. with increasing length of stay in the country. We can then define the emigration hazard
Estimating the piecewise constant hazard model specified in equation (5) returns emigration hazards that are constant within each distinct covariate pattern. The corresponding waiting times, computed by inverse transform sampling, as discussed in Section 3.3.1, are exponentially distributed. The number of events within an interval is Poisson distributed.
48
The use of (piecewise) constant hazard models is convenient because the exponential distribution is memoryless: Having already survived for some time does not affect the remaining expected survival time. At any point in time, the expected waiting time is given by
Figure 3 plots the estimated emigration hazards and the corresponding survival rates for a man who immigrates at the age of 18 and lives in Vienna, for two different country clusters. The upper pane shows the results for a man born in a high-income EU member state in Northern or Western Europe, such as Denmark; the lower pane for a man born in Syria, a country whose emigrants have a long duration of stay and a high number of asylum applications in Austria. The figure shows substantial differences in emigration behaviour, with immigrants from the Northern and Western EU member states experiencing much higher emigration hazards in the first ten years following immigration. With the cohort-component projection model used by Statistics Austria for its national population projections before switching to STATSIM, individuals from both country groups would have been assigned the same, duration-independent emigration rates.

Emigration hazards and survival rates by country cluster and duration of stay for an 18-year-old male who immigrates to Austria and lives in Vienna. Notes: Hazard: Rate at which a person emigrates in a given time interval. Survival: Proportion of individuals who do not emigrate until a given point in time.
In the simulation, each foreign-born person draws a waiting time until emigration based on the estimated hazards. For those born in Austria, a waiting time until emigration is determined by age-, province- and sex-specific probabilities; there is no dependence on the length of stay. The total number of emigrants per projection year is calculated as the sum of individuals for whom an emigration event occurred in the simulation for that year.
The current version of STATSIM exclusively uses administrative data, derived from population registers, as input for its base population and parameters. As the underlying registers contain individual-level IDs, records can be linked across data sets. The data are available at Statistics Austria and cover the resident Austrian population. In principle, STATSIM could also be run with data from a large, representative sample of the population or a synthetic population. For a recent example of a dynamic microsimulation model that uses a synthetic population, see Münnich et al.. 18
For its base population, STATSIM uses data on the population level as of January 1st of the starting year disaggregated by age (in single years), sex, federal province of residence, country of birth (clustered) and duration of stay, which are taken from the Population Statistics. 51 Demographic parameters for the projection of births, deaths and migration are based on data from the Vital Statistics 52 and the Migration Statistics 53 in connection with the Population Statistics. For a documentation of the assumptions on future demographic behaviour made in the 2022 round of population projections, see.54–56
The cluster analysis of origin countries for the international migration module incorporates additional data from the Asylum Statistics 57 and the Register-based Labour Market Statistics. 58 Emigration hazard estimation was performed on individual-level data for 2019, which was a relatively stable and representative year in terms of migration for Austria.
As with all statistics based on administrative data, the migration statistics reflect a reporting reality that may deviate in individual cases from the lived reality of the people under consideration. Migration movements that are not reported to the authorities are not included in the data. On the one hand, this concerns persons residing illegally in Austria who avoid being registered by the authorities. A second group includes people who move abroad without de-registering with the competent registration authority (as required by law). This results in a partial under-recording of emigration due to a lack of de-registration. The registration authorities aim to record these movements retrospectively on an ongoing basis and feed them into the registration database through official de-registrations. This usually takes place in a different reporting year than the actual date of emigration. Nevertheless, there is an ongoing adjustment, which minimises the effect of underestimation over time. The measurement error is further reduced by taking into account the results of the register censuses.
Overall, the registration data from the central population register is of very high quality, as it is filled in directly by the registration offices (registrations and de-registrations of residences), registry offices (births and deaths) and naturalisation authorities (changes of citizenship). According to the Registration Act, the registration or de-registration of a place of residence in Austria must be carried out at the responsible registration office within three days of moving into or out of the accommodation. This information is therefore highly accurate and reliable due to the official character of the documents. In addition, the personal identifier enables the unique assignment and linking of different registration sequences of a person. This also makes it possible to compare individual characteristics (e.g. nationality, country of birth, country of origin or destination) over time, which contributes to a further improvement in data quality. For more information on the Migration Statistics and the quality of the underlying register data, see Statistics Austria 53 .
Results
Between 2001 and 2022, the size of the Austrian population increased by around 1.02 million or 12.7%. Of this increase, only a small percentage (around 1.4%) was due to birth surpluses - the lion's share were gains from migration. 56 In the coming decades, the large birth cohorts of the 1950s and 60s will reach the end of their life span, causing a strong increase in the number of deaths. Unless fertility increases dramatically over this period, from its currently low value with a TFR of 1.32 (2023), natural population change in Austria will be continuously negative throughout the coming decades. 59 In this case, any positive population growth in Austria will be caused by migration surpluses. Figure 4 demonstrates this pattern using the results of Statistics Austria's 2023 population projection, computed with STATSIM. A detailed discussion of the assumptions, data and results of this projection is provided in Slepecki and Pohl. 56 Alongside the main variant, a projection scenario with no internal or international migration was computed. According to this scenario, the size of the Austrian population would decrease by 24.4% until 2080, while according to the main variant, it would increase by 13.1%. These results demonstrate the fundamental impact of migration on population change in Austria and, consequently, the need for a projection model that captures migration behaviour well.

Population change in Austria, 1950-2022 (observed) and 2023-2080 (projected).
To assess the impact of the extended international migration module, as described in Section 3.3.4, on projection accuracy, we compute three retrospective projections by microsimulation, using 2013 as the staring year, and evaluate the results until 2021. The first one implements the cohort-component method and follows Statistics Austria's 2013 population projection. Demographic event rates by age, sex and federal province of residence, as used in the 2013 projection, were converted into waiting times using the inversion method (see Section 3.3.1) and used as inputs. We refer to this model as ‘Cohort-component method, base rates’. The second model extends the first by adding country of birth (native vs. foreign born) as an additional dimension to all event rates. We call this model ‘Cohort-component method, extended rates’. The third model is STATSIM. For fertility, mortality and internal migration, it uses the same input parameters as the second model. For international migration, the extended modelling described in Section 3.3.4, with duration-dependent emigration hazards and detailed clusters for country of birth, is used. Hazards are estimated from 2012 data and the cluster analysis uses data for the years 2010 to 2012. Due to limited data availability for these years, a subset of variables is used to perform the cluster analysis (sex ratio, mean age, standard deviation of the age distribution, share of individuals who stay in Austria for at least six months / one year / two years after immigration, number of university students relative to number of immigrants, applications for asylum or subsidiary protection relative to number of immigrants). Since immigration flows were much higher over the evaluation period than assumed in the 2013 projection, we replace the assumed values with the observations in all three models, which facilitates comparison of the projection results with the actual population and emigration figures. We chose to replicate the 2013 population projection as this allowed us to assess the impact of extending the microsimulation model beyond the cohort-component framework on projection accuracy over the period from 2013 to 2021.
The differences between the models become evident when we contrast projected emigration and population levels with observed values. As shown in Figure 5, the cohort-component model with base rates does not capture the increase in emigration following the high levels of immigration in 2015 and 2016, while the cohort-component model with extended rates and STATSIM both show an increase in emigration. The root mean square error (RMSE) of the projections ranges from 15 773 (cohort-component model with base rates) to 8 975 (cohort-component model with extended rates) and 7 032 (STATSIM).

Projected and observed emigration from Austria 2013- 2021, Cohort-component method vs. STATSIM.
Regarding population size, Figure 6 demonstrates the divergence of projection results from observed values following the years of high immigration (2015/2016). This results in an RMSE of 99 604 for the cohort-component model with base rates, 48 814 for the cohort-component model with extended rates and 17 774 for STATSIM. The figures show that each layer of detail added to the model improves projection accuracy over the evaluation period. Although this result cannot be generalised to imply that increased complexity always enhances projection accuracy, it shows that even (relatively) small model extensions can yield significant improvements if they capture important demographic relationships. Microsimulation models provide the necessary flexibility to account for such interrelations.

Projected and observed population of Austria 2013-2021, Cohort-component method vs. STATSIM.
The cohort-component method is the standard tool for producing population projections in official statistics. It is well documented in the literature, easy to implement and does not require a broad range of input data, making it a valuable technique for practitioners. However, the method is limited in its ability to capture complex demographic processes and interrelations, model diverse populations and generate outcomes for a wide range of individual-level attributes. Dynamic microsimulation presents a solution: This class of models simulates (somewhat) realistic life paths at the individual level, considering their inherent complexity and interdependencies. By modelling demographic behaviour in a more realistic way and producing richer output, microsimulation-based population projections have the potential to be more useful for policy and planning purposes.
Moving from the cohort-component method to a dynamic microsimulation represents a fundamental change in the way a population projection is produced. We opted for a step-wise transition: First, we built a microsimulation model that replicates the results of the multi-state cohort-component method with age, sex, province of residence and native vs. foreign born, as previously used by Statistics Austria for its population projections. Then, as a first step in extending the model beyond the cohort-component framework, we implemented a new module for international migration, which includes length of stay and more detailed information on country of birth as additional variables.
We chose this gradual approach for three reasons: Firstly, replicating the results of cohort-component projections is considered a good model validation for dynamic population-based microsimulations. 26 Secondly, we believe it is easier for the users of Statistics Austria's projections to follow and understand changes in the projection methodology if we introduce them gradually over time. Lastly, it is more convenient for implementation purposes to start with a reasonably simple model and extend it step-wise. This way, even a small team can make the transition from the cohort-component method to microsimulation and develop the model further in line with capacity, i.e. only increasing complexity to an extent that is manageable given available resources. Fortunately, seemingly small model extensions can already have a big impact, as demonstrated by our comparison of STATSIM and cohort-component projection results in Section 5.
Nevertheless, moving from the cohort-component method to microsimulation demands a deeper understanding of model building, advanced statistical programming and data analysis skills as well as more resources (e.g. staff, training) and computing capacities. 4 The inclusion of additional variables also requires more data. Furthermore, while cohort-component projection models can look very similar across different institutions, microsimulation-based projections are usually more tailored to specific use cases and relevant forms of population heterogeneity, making the models less transferable and hence creating a barrier for uptake.
This article reflects the status of STATSIM as it was used to produce the national population projections for Austria in 2022 and 2023. In the future, we plan to gradually develop and extend individual model elements in order to enhance the population projection and produce results for other demographic and socio-economic characteristics. In particular, we plan to expand STATSIM to account for heterogeneity among migrants in fertility and mortality, in addition to emigration, and to extend the model from its demographic core to include modules for education and employment. This will enable us to derive projections for the education level and employment status of the population and its subgroups. Furthermore, we will be able to model demographic processes dependent on individual-level education and employment characteristics, e.g. modelling women's fertility dependent on their education level and employment status.
Footnotes
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
i
If all characteristics are fully interacting in the model parameters, the state space will be equally large in microsimulation and cohort-component models. However, usually we can assume some independence, reducing the size of the state space for microsimulation models [2].
ii
Age in years was grouped in the following way: 0 to 15, 16 to 19, 20 to 25, 26+ years.
iii
Statistics Austria produces a new round of population projections every three years. For each new round, the assumptions regarding future demographic behaviour are assessed and revised. In between rounds, the population projection is updated annually using a new base population but adapting parameters only if necessary.
Appendix
For the 2022 round of population projections
iii
the cluster analysis used to form country groups for the international migration module of STATSIM was performed on country-level data for persons who immigrated to Austria from 2017 to 2019, for all 85 countries with at least 300 immigrants arriving during this period. Countries that were excluded from the analysis due to their small numbers of immigrants in Austria were manually assigned to the established clusters based on geographic proximity. We consider data on (recent) migrant inflows to be more informative for modelling future emigration behaviour than the stock of foreign-born individuals, which is largely made up of long-term residents with a very low propensity to emigrate. Data for the years 2020 and 2021 as well as 2015 and 2016 were excluded, as these years are considered outliers in terms of migration. The former because of the impact of the COVID-19 pandemic, the latter due to the high number of refugees, particularly from Syria and Afghanistan, who migrated to Austria in these years. The following characteristics, aggregated at the country level, were included in the cluster analysis:
Sex ratio Age: Share of individuals below age 14 / between ages 20 and 24 / between ages 25 and 29 Length of stay in Austria: Share of individuals who stay in Austria for at least six months / one year / two years after immigration Share of individuals enrolled in tertiary education Share of individuals in active employment Share of individuals with at least one child under the age of 15 in the family Applications for asylum or subsidiary protection relative to number of immigrants
The choice of variables is based on the assumption that they cover different determinants of emigration behaviour: age, sex, educational enrolment, employment status, family composition, length of stay and asylum and subsidiary protection.
The cluster analysis was performed for three regional groups: EU member states, other countries in Europe, and the rest of the world. EU member states were considered separately due to the right to freedom of movement 60 ; other European countries were analysed independently from the rest of the world due to their geographic proximity to Austria, which impacts migration behaviour. 61
The optimal number of clusters within each region was established by hierarchical clustering. Given these results, a K-Means algorithm was applied to make the final assignment of the countries to the clusters. Through this process, we defined 17 clusters, five of which were in the EU region, four in Europe outside the EU, and eight in the rest of the world.
Data on age, sex and length of stay were derived from the Migration Statistics and averaged over the years 2017 to 2019. The data on immigrants enrolled in tertiary education, those in active employment and those with at least one child under the age of 15 in the family come from the Register-Based Labour Market Statistics, which use October 31st as the reference date. For the cluster analysis, we took the average over the three reference dates. The observations refer to individuals who immigrated between 2017 and 2019, on October 31st of the year in which they immigrated. Immigrants who arrived between November 1st and December 31st were not considered in the formation of these three variables. For asylum and subsidiary protection, we took the number of applications divided by the number of immigrants per country and year. Data were derived from the Asylum Statistics and the Migration Statistics, respectively. As the Asylum Statistics only contain information on citizenship and not country of birth, we took the number of applications by citizenship of a given country as a proxy for the number of applications by birth country.
