Abstract
We introduce a new command,
1 Introduction
Conventional methods to fit fixed-effects dynamic panel regression models require that a researcher observe three consecutive time periods or two pairs of two consecutive time periods. On the other hand, there are some panel datasets with irregularly spaced time intervals that do not satisfy these conventional time-spacing requirements for identification and estimation. For example, personal interviews were conducted in 1966, 1967, 1969, 1971, 1976, 1981, and 1990 for the National Longitudinal Survey (NLS) Original Cohorts: Older Men, and there are neither three consecutive time periods nor two pairs of two consecutive time periods in this list of years. Other examples with irregularly spaced panel data from the United States include Current Population Survey, Early Childhood Longitudinal Survey-K, NLS of Youth 1979, and Panel Survey of Income Dynamics—see section 2 for more details about specific survey periods of these datasets.
Even if panel data exhibit such irregular and unequal time spacing, Sasaki and Xin (2017) show that the fixed-effects dynamic panel regression parameters can be still identified as long as “two pairs of two consecutive time gaps” are available, which generalizes the conventional requirement of “two pairs of two consecutive time periods”. In the NLS Original Cohorts: Older Men, for instance, there is a time gap of 0 between 1966 and 1966, a time gap of 1 between 1966 and 1967, a time gap of 2 between 1967 and 1969, and a time gap of 3 between 1966 and 1969. Thus, there are two pairs, (0, 1) and (2, 3), of consecutive time gaps in this panel dataset. This requirement is also satisfied by the list of the aforementioned panel datasets: Current Population Survey, Early Childhood Longitudinal Survey-K, NLS of Youth 1979, and Panel Survey of Income Dynamics.
In this article, we introduce the
2 Review of the method
This section reviews a part of the method of identification and estimation of fixed-effects dynamic panel regression models under unequal time spacing proposed by Sasaki and Xin (2017). Consider the model
where yit
denotes an observed state variable, xit
denotes an observed covariate, αi
denotes an unobserved individual fixed effect, and εit
denotes an unobserved idiosyncratic shock. A researcher is often interested in the autoregressive parameter γ and the regression parameter β in such a model. The following two examples illustrate concrete regressions in which economists would be interested.
For convenience of writing, we first define a few shorthand and auxiliary notations. Write Ei (·) := E(·| αi ) for the expectation conditional on individual i’s specific fixed effect αi . With this notation, we in turn define the auxiliary random variables
for each time t and time gap τ. Furthermore, we denote their cross-sectional means by Zτ := E(Zi τ ), zτ := E(zi τ ), ζτ := E(ζi τ ), and ζ−τ := E(ζi, −τ ).
Let T be the set of unequally spaced time periods t for which a researcher observes the panel data
We also define the set of gap-associated survey years by
for each time gap τ ∊ T and let T(τ) = ∅ if τ ∉ T. For the NLS Original Cohorts Older Men, for example, personal interviews were conducted in 1966, 1967, 1969, 1971, 1976, 1981, and 1990. In this case, we can write T = {1966, 1967, 1969, 1971, 1976, 1981, 1990}, T = {0, 1, 2, 3, 4, 5, 7, 9, 10, 12, 14, 15, 19, 21, 23, 24}, T(0) = T, T(1) = {1966}, T(2) = {1967, 1969}, T(3) = {1966}, T(4) = {1967}, T(5) = {1966, 1971, 1976}, and so on.
If panel data with unequal spacing have T(1) ≠ ∅, T(Δt) ≠ ∅, and T(Δt + 1) ≠ ∅ for some gap Δt in the set of natural numbers, then we call its spacing structure the “U.S. spacing” (compare Sasaki and Xin [2017, def. 2]). For instance, one can verify that each of the following datasets has this U.S. spacing structure:
See example 2 in Sasaki and Xin (2017) for detailed discussions.
Under the U.S. spacing, Sasaki and Xin [2017, corollary 1 (ii)] show that (γ, β)′ can be identified by
with
provided that | Δ| ≠ 0 holds.
Given the identification result, one can now take the sample counterpart of it to obtain an estimator. The auxiliary random variables, Zτ , zτ , ζτ , and ζ−τ , may be estimated, respectively, by
for any
where
where
While the above procedure focuses on the just-identified case for the sake of clarity, the generic generalized method of moments (GMM) restriction is provided by
where θ = (γ, β),
for all (t, t′, t′′, t′′′) ∊ T(0) × T(1) × T(Δt) × T(Δt + 1). The
where
The variance matrix of the GMM estimator is calculated based on the asymptotic normality
where
See theorem 2 of Sasaki and Xin (2017).
In addition to the U.S. spacing, there is another spacing pattern, called the “U.K.spacing” (Sasaki and Xin 2017, ex. 1). However, the
The
Before forming the moment restrictions presented above, the
3 The xtusreg command
3.1 Syntax
The syntax of the
Here depvar stands for the dependent variable y, and indepvars include independent variables x. Exactly one depvar variable should be included to run the command, while indepvars are optional and may include multiple variables.
3.2 Options
The
3.3 Stored results
4 Simulation studies
In this section, we present the finite sample performance of the
We generate data following the time-spacing pattern of the NLS Original Cohorts: Older Men. First, we consider a simple data-generating process with just a state variable y and without an observed predictor x. Independently generate
for individuals i = 1,…, N, where N = 1000. For the subsequent time periods t = 2,…, 70, we iteratively and autoregressively generate
where we set γ = 0.5 for the autoregressive coefficient. After we have accumulated the whole data [(yit ) : i ∊ {1,…, N}, t ∊ {1,…, 70}] for 70 periods, we drop all the time periods except for t = 1966, 1967, and 1969. This leaves us with [(yit ) : i ∊ {1,…, N}, t ∊ {1966, 1967, 1969}], similarly to the first three survey years for the NLS Original Cohorts: Older Men. The first row of table 1 shows simulation results based on 2,000 Monte Carlo iterations.
Baseline simulation results
The displayed statistics include the bias, standard deviation (SD), root mean squared error (RMSE), and 95% coverage frequency for the parameter, γ. The bias, SD, and RMSE are all small relative to the magnitude of the parameter value γ = 0.5, and the coverage frequencies by the 95% confidence interval are close to the nominal probability of 0.95. We repeat this simulation analysis using extended lists of unequally spaced time periods. The results based on T = {1966, 1967, 1970} are displayed in the second row in table 1. The results based on T = {1966, 1967, 1971} are displayed in the third row in table 1. The results are very similar to those found in the first row, and therefore the same conclusion applies.
Next, we consider a data-generating process with an observed predictor x as well as the state variable y. Independently generate
for individuals i = 1,…, N, where N = 1000. For the subsequent time periods t = 2,…, 70, we iteratively and autoregressively generate
where we set γ = β = ρ = 0.5. Again, we drop all the time periods except for t = 1966, 1967, and 1969, leaving us with [(yit, xit ) : i ∊ {1,…, N}, t ∊ {1966, 1967, 1969}].
The first row of table 2 shows simulation results based on 2,000 Monte Carlo iterations. We repeat this simulation analysis using extended lists of unequally spaced time periods. The results based on T = {1966, 1967, 1970} are displayed in the second row in table 2. The results based on T = {1966, 1967, 1971} are displayed in the third row in table 2.
Simulation results for a model with a covariate
These unequally spaced time periods cannot be handled by conventional commands, such as
5 Illustration of the command
In this section, we illustrate the
In the software package, a subsample of this dataset can be loaded by the following command line:
These data contain six variables:
Having set the panel-data structure, we first run the simple autoregression of y =
The estimate of the autoregressive parameter γ is a significant positive below one value, implying that the log income is positively autocorrelated and follows a stationary process. (Note that income was received in the years {1965, 1966, 1968}, while personal interviews were conducted in years {1966, 1967, 1969}.)
We next include
The inclusion of the control affects the point estimate of the autoregressive parameter γ, but the level of statistical significance is larger than before. Furthermore, the coefficient β of
The dataset has two remaining candidates,
The autoregressive parameter γ is larger for white individuals than for nonwhite individuals, implying that log income is more persistent for white individuals than for nonwhite individuals.
Finally, we also run the fixed-effects dynamic panel regression controlling for
The autoregressive parameter γ is larger for those individuals with 12 years of education or higher, implying that log income is more persistent for this group of individuals.
Existing commands could not have produced these results, because of the irregular time spacing of the NLS Original Cohorts: Older Men. The new command,
6 Conclusion
Conventional methods to fit fixed-effects dynamic panel regression models require an observation of three consecutive time periods or two pairs of two consecutive time periods. Sasaki and Xin (2017) show that the availability of “two pairs of two consecutive time gaps”, which generalizes the conventional requirement of “two pairs of two consecutive time periods”, suffices for the identification of the model parameters. In this article, we introduced the
Finally, we discussed a limitation of the
7 Programs and supplemental materials
Supplemental Material, sj-zip-1-stj-10.1177_1536867X221124567 - xtusreg: Software for dynamic panel regression under irregular time spacing
Supplemental Material, sj-zip-1-stj-10.1177_1536867X221124567 for xtusreg: Software for dynamic panel regression under irregular time spacing by Yuya Sasaki and Yi Xin in The Stata Journal
Footnotes
7 Programs and supplemental materials
To install a snapshot of the corresponding software files as they existed at the time of publication of this article, type
