Abstract
In spatial epidemiology studies, the effects of covariates on adverse health outcomes could vary over space and time so examining the spatio-temporally varying effects is useful. In particular, the association between covariates and health outcomes could have locally different temporal patterns. In this article, we develop a Bayesian spatio-temporal latent model to identify spatial clusters in each of which covariate effects have homogeneous temporal patterns as well as estimate heterogeneous temporal effects of covariates depending on spatial groups. We compare the proposed model to several alternative models to assess the performance of the proposed model in terms of a range of model assessment measures. Low birth weight incidence data in Georgia for the years 1997–2006 are used.
1 Introduction
Numerous epidemiology studies have mainly focused on the association between risk factors and health outcomes such as child-related diseases. Health data and covariate data are often collected over space and time and they have space-time variation. So, the association between covariates and health outcomes may vary spatio-temporally. In particular, the relations in spatial health data could be different depending on areas within the spatial domain. They could also have locally different temporal patterns in the spatio-temporal domain. However, the commonly used spatial model assumes that covariate effects on health outcomes are constant within their entire space and time domains, although relative risk is constructed with a function of space-time random effects. 1 – 4 Thus, investigating the spatio-temporal association between covariates and health outcomes is useful and important.
In this article, the focus is on identifying the spatial groups where covariate effects have different temporal patterns depending on groups as well as estimating their temporal profiles varying locally. In analyzing small area health data, covariates may have different temporal effects on disease outcomes, depending on geographical areas (e.g. urban/rural areas). In some cases, researchers often consider dynamic models to capture spatio-temporal variation on the coefficients in linear models 5 . But, these approaches have difficulty to estimate the locally heterogeneous temporal profiles of the covariate effects. Recently, Lawson et al. 6 proposed a Bayesian spatio-temporal mixture model for analyzing the spatio-temporal health data to estimate temporal patterns varying across space in relative risks. They only focused on understanding the heterogeneous temporal patterns of relative risks not covariate effects. So, the development of statistical modeling is needed to find the spatial groups that are determined by distinct temporal patterns of the covariate effects and then estimate the locally varying effects.
To investigate the locally varying temporal behavior of covariate effects, we develop a new Bayesian latent model with space-time varying coefficients. The proposed model allows for the heterogeneous temporal patterns of coefficients depending on areas within the study domain. In each spatial group, a coefficient has a temporal profile which is different from the profiles of the coefficients in other groups. Since the temporal patterns of the covariate effects within neighboring areas may be different and the patterns in disconnected areas may be same, the proposed model does not require the groups to be spatially adjacent, which is a general assumption. Instead, we model the group memberships of coefficients using spatially dependent weight structures to consider spatial variation in grouping patterns. A Bayesian hierarchical approach is also employed to model locally temporal profiles of the coefficients as well as space clustering in the coefficients simultaneously. Thus, the novelty of the proposed model is flexibility and practicality in analyzing space-time health data. To our best knowledge, this work is the first attempt to consider the spatial groups where the covariate effects have distinct temporal profiles within a statistical model and then better understand the locally varying temporal effects of covariates.
In this article, we assume that the number of spatial groups is finite and can be estimated. However, the determination of the number of spatial groups is one difficulty in clustering approaches. To avoid the specification of the number of groups, Dirichlet process mixture modeling or the use of entry parameters can be considered. 7 – 9 However, these approaches have a computational burden because the maximum number of groups considered in the model should be assumed to be very large and dramatically increased with the size of the spatial domain. In our proposed modeling, we overcome this problem by using several model assessment criteria. We initially consider the range of the number of spatial groups based on the spatio-temporal variation of the data of interest and the size of the spatial domain. The proposed model with different number of groups is fitted and the best number of spatial groups in the proposed model is determined by using various model diagnostics. This approach is especially appropriate for small area health data because the number of spatial groups in small area data would be finite in general.
To assess the performance of the proposed model, we conduct the comparison of our proposed model to several alternative models that do not identify the heterogeneous behavior of coefficients depending on areas. The competing models include constant coefficients over space and time, temporally varying coefficients over space, spatially varying coefficients over time, and globally space-time varying coefficients. A Bayesian model choice criterion 10 and prediction measures 11 are considered to examine the performance of the models considered. For the comparison, a county-level low birth weight (LBW) incidence data set from the state of Georgia is used. We also explore the results of the covariate effects on LBW from the proposed model.
The outline of this article is as follows. In Section 2, we describe the LBW incidence data used in this research. Section 3 proposes our spatio-temporal latent model, and Section 4 explores various model comparison methods. In Section 5, we apply our spatio-temporal model and alternative models to the LBW data, find the best model using model assessment tools, and provide the findings. Conclusions and potential future work are provided in the last Section 6.
2 The data
We used county-level LBW (infant birth weight less than 2500 g) incidence data in Georgia for the year 1997 to 2006. The counts of LBW were collected from the state health information system OASIS (Georgia Division of Public Health, http://oasis.state.ga.us/).. The number of counties in Georgia is 159 and the number of years is 10. As socioeconomic covariates of LBW, 12 we consider the county-level population density (defined as population divided by total land area in square miles), the proportion of black people, median household income, and unemployment rate, which were obtained from the US census and the US Bureau of Labor Statistics. In addition, we use aggregate data based on birth certificates for the other socio-demographic and behavioral risk factors during pregnancy related with mothers. The proportion of mothers with less than 12th grade education, the proportion of mothers smoking during pregnancy, and the proportion of mothers with ‘Inadequate’ value from the Kotelchuck Index 13 (IKI value) are considered.
Figure 1 displays the maps of standardized incidence ratios for LBW births where the standardized incidence ratio is defined as the number of LBW births divided by the expected counts computed using the internal standardization method.
14
Here, the spatio-temporal variation of standardized incidence ratios can be seen. For example, the LBW standardized incidence ratios in some central areas of Georgia are higher than other areas over years. In south-east areas, the standardized incidence ratios have temporal variation, showing that the ratios from 2001 to 2003 are overall lower than those for the other years. We also found the spatio-temporal variation of the covariates considered, so the effects of covariates on LBW in Georgia may vary across space and time. In addition, the spatial domain (Georgia state) is a small area so it is reasonable to assume that the number of spatial groups is finite and small. Thus, we apply the proposed model and the competing models to the LBW data and assess the performance of the models.
The standardized incidence maps of county-level LBW data in Georgia for each year. LBW: low birth weight.
3 Statistical model
Let {Y
it
: i = 1, … , I, and t = 1, … , T} be the LBW count data collected for county i at time point t, where I = 159 and T = 10 and let E
it
be the expected count. Conditional on E
it
, we model Y
it
as a standard Poisson distribution
In this study, we assume that there are spatial clusters in which each has a set of homogeneous time-varying coefficients. Thus, covariate effects on LBW change over spatial clusters and they within a spatial cluster have own temporal pattern. To satisfy this assumption, we consider
To construct spatial clusters, we model Z(i) as
The coefficient vector
From the previous equations, we can easily express the likelihood of the observed counts
In Bayesian mixture modeling, component identifiability problems can emerge since the likelihood is not variant with respect to the permutation of the components labels, which was found in Stephens. 17 Recently, Choi et al. 9 examined when label switching problems appeared in spatio-temporal mixture modeling and found out that components could switch labels during MCMC simulation if multiple chains were used. We thus use a single chain run in this analysis to avoid this label switching problem.
4 Model comparison
As a Bayesian model selection method, the deviance information criterion (DIC) of Spiegelhalter et al.
18
is widely used. DIC forms the goodness of fit as well as the complexity of the Bayesian model. It is defined as
In this article, we also consider a comparison method through the conditional predictive ordinate (CPO).19,20 This CPO is a cross-validation measure and obtained using the marginal posterior predictive density so the CPO for county i at time t is defined as
Along with the DIC3 and the MPL, the mean square prediction error (MSPE) is finally considered for the comparison of models in terms of prediction performance
5 Data analysis and results
The results reported in this section are based on posterior samples of 70,000 iterations per MCMC, with 20,000 burn-in iterations. We also keep every 10th sample to avoid the high autocorrelations of the some quantities. Thus, 5000 final samples are used for the estimation of the parameters. To ensure MCMC convergence, several diagnostics such as the Geweke convergence diagnostic, 22 autocorrelation functions, and trace plots are used. Both R (http://www.r-project.org) and WinBUGS (http://www.mrc-bsu.cam.ac.uk/bugs) are used to implement the models considered in this article.
We now apply our proposed model to the county-level low birth weight incidence data as described in Section 2. In order to select the best number of spatial groups in the model, the model with a range of the number of groups is fitted. The number of spatial clusters considered here is between 2 and 10, where the number of clusters is assumed to be at least two. The proposed model with the best number of groups is determined using several model comparison measures described in the previous section. In Figure 2, the plots of DIC3, MPL and MSPE values for the proposed model with the different number of spatial groups are displayed. Overall, we find that DIC3 and MSPE values decrease and MPL value dramatically increases as the number of spatial clusters (G) increases up to 8. DIC3 and MSPE values for the proposed model with 7–9 spatial groups are quite similar and small even though the pattern of MPL values for the model with 7–9 groups is a little bit different. In this analysis, we can see that the model with 10 groups (the maximum number of groups considered) does not provide the best improvement among the model with the different number of groups considered in terms of these model selection measures. MPL measure indicates that the model with 8 spatial groups has the largest value among the models considered. In addition, the proposed model with 7 or 8 groups has similar small DIC3 and MSPE values so it is reasonable that the model with 8 spatial groups provides the most adequate fit.
The selection of the number of groups in the proposed model using DIC3, MPL and MSPE. DIC: deviance information criterion; MPL: marginal predictive-likelihood; MSPE: mean square prediction error.
We also compare the proposed model with four competing models in our analysis. The models under consideration are as follows:
Model comparison using DIC3, MPL and MSPE for LBW data
DIC: deviance information criterion; MPL: marginal predictive-likelihood; MSPE: mean square prediction error; LBW: low birth weight.
The map of the spatial cluster indicator including 8 groups and the histogram of the number of counties by group are presented in Figure 3. The fifth and seventh groups have the largest number of counties in Georgia, which is 39 counties (24.5%) for each spatial group. In addition, the downtown of Atlanta is assigned to the seventh group. The eighth spatial group has the third largest number of counties and it includes 27 counties (17.0%). Many counties in the Atlanta suburbs are especially allocated to the eighth group. On the other hand, 1 county (Charlton county) is only assigned to the first group and six counties located in east areas of Georgia are assigned to the sixth group. The remaining groups (groups 2, 3 and 4) include between 12 and 20 counties.
The map of 8 spatial groups (left) and the histogram of the number of counties by group (right).
We find that the temporal profiles of the covariate effects on LBW vary across spatial clusters. For example, the coefficient corresponding to the proportion of mothers with low education in the fourth spatial group dramatically increases over time as presented in Figure 4. This coefficient in group 3 is also a little increasing over time. Even though the coefficient in groups 7 and 8 has a stable temporal trend, the estimated coefficient in group 7 is smaller than that in group 8 over time. This indicates that some of southern-east and southern-west areas (groups 3 and 4) have increasing temporal effects of the proportion of low educated mothers over time while the downtown of Atlanta and some of the Atlanta suburbs have stable temporal effects. It is also found that the estimates of the coefficient are positive even though they are not significant. They thus suggest that a higher proportion of low educated mothers is associated with increased risk of the LBW incidence and the effects of maternal low education level in rural areas are higher than those in urban areas. These findings are consistent with previous analyses.23,24
Temporal plots of the coefficient corresponding to the proportion of mothers with low education for the selected spatial groups (Group IDs 3, 4, 7 and 8). The solid lines indicate the posterior mean and the dotted lines indicate the 95% confidence intervals.
To evaluate the spatial prediction performance of the proposed model, we conduct calibration analysis. First, we randomly selected 16 counties (about 10% of the data) and fitted the proposed model 16 times by removing all observations from the selected county. For the second analysis, we removed the observations from all 16 counties simultaneously at a single time point, repeating this analysis for each of the 10 observation times. For each case, the 160 estimated LBW counts Ŷ
it
are obtained and compared with the observed counts Y
it
. Figure 5 presents the calibration plots for the LBW. The percentage of the observations that are outside the 95% intervals is 3.85% for the first calibration design (Figure 5(a)) and 3.12% for the second one (Figure 5(b)). We thus concluded our proposed model in this application performed well in terms of the spatial forecasting ability.
Model diagnostics for LBW data. The dotted lines show the 95% prediction intervals. (a) First design, (b) Second design. LBW: Low birth weight.
To explore whether the random walk process assumption in the coefficients considered is reasonable in this data set, we refit the proposed model with 8 spatial groups where the coefficients have an autoregressive distribution with order 1, AR(1), (
6 Conclusion
In this article, we introduced a novel and flexible Bayesian latent model for space-time health data that accounts for the spatio-temporally varying coefficients. The proposed model captures the locally varying temporal behavior of covariate effects. We show, using low birth weight incidence data, that the latent model with spatio-temporally varying coefficients proposed in this paper provides a substantial improvement over other competing models that do not have such properties in terms of a range of model comparison measures.
The spatio-temporal latent model introduced here is the first step to illustrate the locally different temporal profiles of coefficients using a hierarchical framework. However, our approach has some limitations. In the real data analysis, the estimated number of groups in the proposed model is quite large (8). It would be because the model assumed that the spatial groups are fixed across covariates. However, the spatial groups could vary with different covariates. If we would consider different number of spatial groups depending on covariates, then we would expect that the model would produce a smaller number of groups. Thus, the spatial grouping structure could be extended to the different spatial groups depending on covariates in the future. We also found that covariates in some spatial groups have little effects on LBW. So, the subset of covariates that have significant effects on health data could vary across spatial groups. As future work, we will consider the development of spatial variable selection approaches within a spatial latent model framework, which allows the subset of important covariates to vary across spatial groups. This research is particularly useful and valuable in spatial health effects studies because covariate effects are significantly different depending on areas as well as the subset of covariates can be identified within a subset of geographical areas.
Footnotes
Funding
This research was supported by NIH Grant R21 R21HL088654-01A2.
