Abstract
Marketing research relies on individual-level estimates to understand the rich heterogeneity of consumers, firms, and products. While much of the literature focuses on capturing static cross-sectional heterogeneity, little research has been done on modeling dynamic heterogeneity, or the heterogeneous evolution of individual-level model parameters. In this work, the authors propose a novel framework for capturing the dynamics of heterogeneity, using individual-level, latent, Bayesian nonparametric Gaussian processes. Similar to standard heterogeneity specifications, this Gaussian process dynamic heterogeneity (GPDH) specification models individual-level parameters as flexible variations around population-level trends, allowing for sharing of statistical information both across individuals and within individuals over time. This hierarchical structure provides precise individual-level insights regarding parameter dynamics. The authors show that GPDH nests existing heterogeneity specifications and that not flexibly capturing individual-level dynamics may result in biased parameter estimates. Substantively, they apply GPDH to understand preference dynamics and to model the evolution of online reviews. Across both applications, they find robust evidence of dynamic heterogeneity and illustrate GPDH’s rich managerial insights, with implications for targeting, pricing, and market structure analysis.
Keywords
The modeling of dynamic phenomena is central to marketing research. Marketers are interested in understanding the evolution of consumer perceptions, preferences, and response sensitivities, as well as the success and failure of different brands over time. Marketing decisions that focus on temporal consequences of marketing actions necessarily rely on empirical models of such marketing dynamics (Naik 2015; Pauwels and Hanssens 2007; Xie et al. 1997). Often, these dynamics are heterogeneous across individual units. We use the term “individuals” to broadly refer to units over which the heterogeneity is defined, examples of which include consumers, brands, and products. For example, the pattern of evolution of preferences could vary across customers because of how they are differentially affected by economic conditions such as recessions. Similarly, how market perceptions evolve could vary across brands because of competitive activity. In such situations, the interest is in the market-level evolution of preferences, as well as in the individual-level trajectories that may differ from each other and from how the market is evolving on average.
In this article, we develop a modeling framework for representing such dynamic heterogeneity. Dynamic heterogeneity characterizes situations where individual-level model parameters evolve over time according to a stochastic process. More specifically, we allow individual-level parameters to evolve flexibly in a fashion that does not force them to exactly mimic the dynamic evolution of the population mean. We do this by allowing the individual deviations from the population means to vary over time. At a given point in time, the collection of individual-level parameters forms a distribution of cross-sectional heterogeneity. The evolution of these individual-level parameters therefore results in a time-varying population distribution, in which the relative positions of individuals change over time. We illustrate the concept of dynamic heterogeneity in the top row of Figure 1.

A synthetic example of dynamic heterogeneity.
While marketing researchers have modeled many different forms of heterogeneity (DeSarbo et al. 1997), most of the literature focuses on the variation in preferences across individuals. Variation within individuals over time has been relatively understudied. Modeling this within-individual variation has important managerial implications for understanding changes in markets over time, and for developing dynamic segmentation and targeting strategies. In addition, just as ignoring cross-sectional heterogeneity can induce estimation bias, not accounting for parameter evolution can also distort inferences about elasticities or response sensitivities and misinform managerial actions.
Several marketing studies have used models that include time-varying individual parameters. Examples include Kim, Menzefricke, and Feinberg (2005), Liechty, Fong, and DeSarbo (2005), Sriram, Chintagunta, and Neelamegham (2006), Lachaab et al. (2006), and Guhl et al. (2018). These studies have used different specifications to capture the evolution of parameters. Kim, Menzefricke, and Feinberg (2005), for example, use a vector autoregressive model to represent the evolution of the population mean, while Sriram, Chintagunta, and Neelamegham (2006) employ a dynamic linear model, and Guhl et al. (2018) rely on penalized splines. Crucially, while all these works use a dynamic model to capture how parameters evolve on average, each imposes a static heterogeneity assumption: conditional on a time-varying mean model
We propose a new methodological framework for modeling dynamic heterogeneity in hierarchical models. Drawing on the literature on Bayesian nonparametric models in statistics and machine learning (Rasmussen and Williams 2005), we develop a novel Gaussian process dynamic heterogeneity (GPDH) specification that characterizes heterogeneity over time-varying latent variables using individual-level random functions of time. These functions are estimated using Gaussian processes (GPs) that are centered around a common mean model. This model captures population-level dynamics and is itself inferred from the data. The GPDH specification is a dynamic analog to static random coefficient specifications, where the mean model plays the role of the population trend and the individual-level functions capture time-varying heterogeneity around this trend. Similar to traditional heterogeneity specifications, our proposed dynamic heterogeneity specification allows for (1) the sharing of statistical information across individuals by shrinking their trajectories toward a common mean trajectory, (2) the sharing of statistical information within individuals across time periods (i.e., intra-individual smoothing), (3) flexible intertemporal evolution, and (4) a principled probabilistic mechanism for projecting the evolution of individual and mean trajectories into other time periods. An important feature of our GPDH specification is that it can be used with any mean model, allowing the researcher to incorporate prior expectations, theory, or even specific drivers of dynamics. Our use of GPs to nonparametrically represent individual-level latent functions that are shrunk toward a common dynamic population model is a novel contribution to the econometric, marketing, and machine learning literatures.
Capturing dynamic heterogeneity yields many benefits. It can generate insights about patterns of individual-level evolution. For example, as we show in our two applications, identifying individuals whose parameters shifted from one extreme of the population to the other, or who moved from being in the extremes of the distribution to the center or vice versa, can enhance managerial and substantive understanding, and managers can leverage this insight for targeted marketing. However, the importance of capturing dynamic heterogeneity goes beyond such individual-level insights. Statistically, if dynamic heterogeneity is present but static heterogeneity is assumed, as is commonly done, we could obtain misleading estimates about both the population-level mean and the extent of heterogeneity, both of which can negatively affect targeting decisions. This can be true even if the correct functional form for the population mean model is used, as we illustrate through simulations.
In this article, we present two applications of GP dynamic heterogeneity. The first and most extensive is in a choice-modeling context, similar to our motivating examples, where GPDH is used to represent time-varying consumer preferences for consumer packaged goods (CPGs), over a span of time that includes the Great Recession. Using both simulated and real purchasing data, we show that GPDH yields more accurate and statistically efficient population and individual-level estimates of preference evolution. On data from six CPG categories, GPDH outperforms static heterogeneity specifications in both fit and forecasting tasks, across a wide array of performance metrics. At the same time, GPDH also uncovers individual-level patterns that can be used to characterize and target customers and to study individual-level responses to economic shocks. More specifically, we find both simulated and empirical evidence of an attenuation bias in estimating population-level parameters when assuming static instead of dynamic heterogeneity around a dynamic mean model. Moreover, we find that, across all categories, estimated individual-level elasticities are notably higher when estimated with dynamic versus static heterogeneity. These biases, and the ability to predict individual-level dynamics, can directly affect category managers’ decision making. Finally, the individual-level dynamics uncover cross-category differences in response to the recession: while there are obvious aggregate-level changes in price sensitivity in many categories, GPDH also uncovers category-level differences in the degree of individual-level response to the Great Recession.
Apart from the modeling of preferences, our specification can be adapted to multiple settings. In our second application, we focus on an entirely different substantive context: the modeling of product reviews. We develop a novel, GPDH-based dynamic topic model to summarize relevant topics that are discussed in customer reviews for different brands of tablet computers. Particularly, our GPDH topic model captures how the content of reviews for individual products evolves relative to aggregate patterns. Empirically, we show how these product-level topic trajectories give insights about the dynamics of market structure in the tablet computer market. Such a granular set of results is not obtainable via aggregate models of market dynamics.
The rest of the article is structured as follows: We first give an overview of the needed methodological background before introducing our GPDH framework. We then discuss our two applications: choice modeling and topic modeling. Finally, we highlight other GPDH applications, describe limitations of the paper, and suggest future research directions.
Methodological Background
Literature
Gaussian processes are Bayesian nonparametric models that are popular in statistics and computer science (O’Hagan and Kingman 1978; Rasmussen and Williams 2005; Williams and Barber 1998) for flexibly modeling temporal and spatial phenomena. Marketing researchers have used Bayesian nonparametrics to represent heterogeneity in static model parameters via Dirichlet process priors (Ansari and Iyengar 2006; Ansari and Mela 2003; Braun and Bonfrer 2011; Braun et al. 2006; Kim, Menzefricke, and Feinberg 2004; Li and Ansari 2014). While Dirichlet processes are most commonly used to model uncertainty over probability distributions, GPs are most commonly used to model uncertainty over spaces of continuous functions. 1
Dew and Ansari (2018) used GPs in a marketing application to decompose variation in purchase rates in a dynamic customer base analysis setting. In their application, GPs were used to represent a mean model of spending rates, but individual-level variation around that mean model was still assumed to be static. Gaussian processes are also related to kriging methods used in Bronnenberg and Sismeiro (2002) for predicting demand across markets. In the context of choice models, Girolami and Rogers (2006) use GPs to model the utility functions of multinomial probit models in a nondynamic and nonheterogeneous context. Finally, our work is also closely related to the marketing literature that models individual-level dynamics (Ansari and Iyengar 2006; Guhl et al. 2018; Khan, Lewis, and Singh 2009; Lachaab et al. 2006; Liechty, Fong, and DeSarbo 2005; Sriram, Chintagunta, and Neelamegham 2006) via FO specifications. In this article, we show how GPDH offers a more flexible and precise alternative than those restricted forms of heterogeneity. Next, we briefly describe GPs. 2
Gaussian Processes
A GP is a stochastic process
where
The mean and the kernel determine the nature of the functions that a GP prior generates. Informally, the mean function encodes the expected location of the functions, whereas the kernel encodes function properties, such as smoothness, amplitude, and differentiability. Much of the GP literature assumes a constant mean function to reflect a lack of prior knowledge about the shapes of the unknown functions, and the kernel serves as the main source of model specification. Many different kernels have been proposed in the GP literature. In theory, the kernel can be any function
Gaussian Process Dynamic Heterogeneity
We now introduce our GPDH specification in the context of a general hierarchical nonlinear modeling framework specified in multiple stages. The first stage models the individual-level data in terms of individual-specific latent functions of time, the second stage specifies how these latent functions vary across individuals according to a GP that is characterized by a mean model and a covariance kernel, and the third stage specifies priors over any invariant parameters in the individual-level model and the hyperparameters for the mean model and the heterogeneity specification.
Stage 1: Individual-Level Model
Suppose that the data
Stage 2: Heterogeneity Specification
The key conceptual innovation of our framework is considering the time-varying individual-level parameters,
The mean function of this GP is a dynamic population model
Mean model
The focus of this work is on capturing dynamic heterogeneity around a focal model, and we thus assume that the researcher has a specific mean model in mind. The marketing literature on dynamic modeling includes several examples, such as state space models (e.g., dynamic linear models that are typically estimated via the Kalman filter in simpler settings), traditional time series models such as autoregressive moving average (ARMA; Box et al. 2015), and parametric models capturing a specific dynamic phenomenon, as in latent force models within the machine learning literature (Alvarez, David, and Neil 2013) or models of advertising dynamics in marketing (Naik, Mantrala, and Sawyer 1998). The mean model could also be another GP. We use different examples in our applications. Which specification is appropriate depends on the modeling context. Again, our goal is not to compare mean models but to illustrate their use in understanding dynamic heterogeneity. Provided that an appropriate (or sufficiently flexible) mean model is used, we have found that the bigger gain in performance comes from using dynamic versus static heterogeneity, rather than the choice of mean model.
Kernel choice
The kernel captures the properties of the dynamic heterogeneity. In this work, we use the rich class of Matérn kernels, which has a general form given by
where
Previous work has shown that the Matérn kernel hyperparameters cannot all be consistently estimated, and, in particular, ν cannot be separately identified from κ (Kaufman and Shaby 2013; Zhang 2004). Hence, ν is typically fixed to a value that reflects the supposed smoothness of the underlying process. Moreover, when the degree is fixed to a half integer (
Fixing ν to a half integer thus makes kernel estimation more tractable. This is especially important when inference methods rely on gradients that involve the kernel function, as derivatives of the Bessel function can be computationally intensive. Furthermore, when the degree
We use the Matérn kernel class in this work for several reasons. First, it is easier to control the smoothness of the function draws from this kernel such that momentary temporal fluctuations can be captured while still representing the underlying smoothness of the process. This is especially suitable for the preference data in Application 1. Second, this class nests the squared exponential kernel—the typical workhorse of the GP literature and used by Dew and Ansari (2018)—as a limiting case. Third, as we describe in the next section, the GPDH specification with the Matérn kernel nests more common heterogeneity specifications as special cases. Finally, Matérn kernels allow the use of complexity-penalizing priors, which facilitates fully Bayesian inference in a principled manner.
Link with static heterogeneity specifications
With this kernel specification, GPDH nests the static (FO) heterogeneity specification as a special case. Mathematically, the FO model assumes

Examples drawing from a GPDH model with a fixed mean function.
Stage 3: Hyperpriors
We employ a fully Bayesian strategy for estimating the GPDH hyperparameters. In particular, we leverage the penalized complexity (PC) prior for Matérn Gaussian random fields introduced by Fuglstad et al. (2018). The PC prior is a weakly informative prior, based on the idea of penalizing the complexity induced by the kernel hyperparameters in the resultant GP. Complexity in classical GP models refers to functions with high amplitude (large η) and small length-scale (small ρ, equivalent to large inverse length-scale, κ). In GPDH, these hyperparameters have distinct meanings: the individual-level amplitude governs the degree of inter-individual shrinkage, while the inverse length-scale captures the degree of individual-level dynamics. Thus, by penalizing high amplitudes and high inverse length-scales, the PC prior encourages shrinkage across individuals, and places substantial prior mass on the nested FO model. The density of the PC prior is
Despite the nonintuitive functional form, another advantage of this prior is that the parameters
In our work, we fix
Estimation
Given the generality of our framework, the details of the estimation procedure for a hierarchical model that uses GP dynamic heterogeneity depend on the specific individual-level model used in the first stage. We discuss our application-specific strategies in the following sections. As a general point, several different inferential strategies have been proposed in the GP literature. These include the use of Laplace approximations, variational Bayesian methods, and expectation propagation methods (Girolami and Rogers 2006; Rasmussen and Williams 2005). Often, approximate inference techniques are used with GPs to overcome the computational complexity in estimating the function values and the hyperparameters of a GP when T, the number of time periods, is large. In our applications, as the number of time periods is not large, we use Markov chain Monte Carlo (MCMC) methods for exact inference. Filippone, Zhong, and Girolami (2013) and Filippone and Girolami (2014) perform a comparative evaluation of different MCMC estimation strategies for GP models. In particular, we use the no-U-turn sampler (NUTS) variant of Hamiltonian Monte Carlo (HMC; Hoffman and Gelman 2014). We have found in our GPDH applications that it is important to jointly sample both the function values and the hyperparameters in one go, as the strong dependency between these sets of parameters makes HMC-within-Gibbs strategies slow to converge.
Distinctions from Previous Work
We reiterate two important features of our approach that make it distinct from previous work. The first is that in our specification, the GPs are used to estimate individual-level functions, which is distinct from using GPs to estimate mean dynamics, as in the work of Dew and Ansari (2018). Second, while recent work by Yang et al. (2016) appears similar to ours in the use of collections of GPs, they model observed variables using GPs. In contrast, we model latent individual-level model parameters via GPs. Because the quantities of interest in our work are latent, we must impose more restrictions than Yang et al. (2016) impose on the nature of the covariance. Specifically, we assume a parametric form for the covariance kernel, which allows us to estimate the model without needing to directly observe the quantity of interest. This assumption also lets us mathematically link the GPDH method to existing heterogeneity specifications as special subcases of our specification.
Application 1: Dynamic Preference Heterogeneity
We now apply our modeling framework to study the evolution of individual-level preferences over time in a multinomial logit choice model. We first estimate the model on synthetic data to illustrate the relative merits of GPDH and the potential pitfalls of not capturing dynamic heterogeneity. We then shift our focus to real data of grocery store purchasing during the Great Recession.
The GPDH Multinomial Logit Model
We consider discrete choice data
such that consumers choose the alternative with the highest utility. For identification, we normalize the intercept of the brand with the highest market share to zero. As we order brands by market share, such normalization effectively forces Brand 1’s intercept to zero. This yields the standard softmax specification for the logit choice probabilites in terms of the individual-level time-varying intercepts and sensitivities in
For both the simulations and the real data, we fix
Mean models
Our emphasis is on modeling the evolution of heterogeneity around a given mean model. Therefore, and to illustrate the flexibility of GPDH, we test four different mean models in this application, corresponding to four common specifications in the literature: Random walk (RW) state space: The RW is the simplest linear state space model that is used in the Kalman filtering literature. Our implementation is given by Gaussian process: As in the work of Dew and Ansari (2018), we can assume a GP as the population model: We assume a constant mean Autoregressive moving average time series: Time series models are especially common in econometric applications and can easily be incorporated into our GPDH framework. We test an ARMA(1) mean-model specification, given by Parametric: A theory-driven parametric model can also serve as the mean model. In this case, one interesting question is the degree to which the Great Recession is associated with changes in consumers’ preference parameters. Thus, to illustrate how a parametric model could be used in conjunction with GPDH, we use a mean function given by the probability density function of a generalized inverse gamma distribution: with
For each of these, we subsequently denote the collection of parameters of the mean model generically as α. Note that α varies across different mean models.
Extensions
There are many possible extensions and alternatives to the utility and mean-model specifications that can incorporate other potentially desirable features alongside dynamic heterogeneity, depending on the available data and choice context. For instance, if the researcher has access to a set of potential drivers of shifts in preferences, such as individual-level events like job loss or changes in income, or market-level events like an indicator for the Great Recession, these can be incorporated directly into the mean model. Specifically, with these drivers denoted generically as
where
A second important consideration in many choice modeling contexts is endogeneity, particularly price endogeneity. Although in this work we focus only on the modeling of heterogeneity, the GPDH specification can be used in conjunction with methods for controlling for endogeneity. For instance, in the case of price endogeneity, the two-stage control function method of Petrin and Train (2010) or the semiparametric approach of Li and Ansari (2014) could be seamlessly incorporated into the utility specification in Equation 7, together with an additional equation for the price-setting process.
Estimation
We estimate all variants of our GPDH logit model via HMC, using the NUTS algorithm. Specifically, we jointly sample all model parameters, including the individual-level functions, the shared mean function, and the hyperparameters. For the parameters of the mean model, we use weakly informative priors. The joint density for the full model is given by
where
Simulations
In this section, we briefly describe a simulation exercise that illustrates the benefits of modeling dynamic heterogeneity. In the Web Appendix, we include additional simulations to help understand the shrinkage properties and the computational complexity of GPDH.
To understand the benefits of capturing dynamic heterogeneity with the GPDH model and the potential limitations of competing approaches, we simulate multiple sets of choice data from the GPDH multinomial logit with a GP mean model. We then estimate the following three models on each of the data sets: (1) the true model (GPDH logit with GP mean); (2) an FO model that uses the GP mean model, but with static heterogeneity; and (3) an independent periods (IP) mixed logit specification that estimates a mixed logit model in each period, with only the variance of the random coefficients shared across periods, which therefore does not directly allow for within-individual shrinkage across time. By simulating data with GPDH, we ensure the presence of dynamic heterogeneity. Moreover, by assuming a GP mean as the true data-generating process, we nest both the FO and IP specifications as limiting cases. 8
Two key results emerge from these choice model simulations. First, by sharing information both within and across individuals, GPDH yields highly efficient estimates, relative to models that assume independence across time periods. By efficient, we mean small credible intervals, while still recovering the true curve. We illustrate this in Figure 3, which shows examples of true individual-level curves and their recovery by the three specifications. As expected, GPDH correctly recovers the true curves with a reasonable amount of precision, as shown by the 95% credible intervals, relative to the curves recovered by the IP model. Under IP, there is no intertemporal sharing of information, leading to estimates that are jagged and with much wider credible intervals. Finally, under FO, the recovered curves are simply wrong: since the FO model assumes that individuals are always at a fixed distance from the mean trajectory, the interesting patterns of individual-level variation are missed.

Illustrative examples of true individual-level curves and their recovery.
The second key result is that, if dynamic heterogeneity exists in the data but static heterogeneity is assumed as in the FO model, the population-level estimates under FO are biased toward zero. This is the case even when the true data-generating mean model is used in the FO model, as it is in our simulations. In Figure 4, Panel A, we illustrate this bias for a single simulated data set, plotting the recovered marginal distribution of the point estimates of the coefficients at different points in time. We can see that the posterior median as recovered by FO is always biased toward zero, and that the estimated distribution of effects has substantially less variation than the truth. In Figure 4, Panel B, we show the same result, but from 172 repeated simulations, where we varied the η parameter of the true GPDH data-generating process. For each simulated data set, we again estimated both GPDH and FO heterogeneity specifications around the same (true) mean model. Then, we computed the mean absolute percentage error in recovering the true population mean. We see that the error is higher in the FO model and increases with η, which represents the magnitude of dynamic heterogeneity in the data-generating process. Taken together, these simulations suggest that the popular approach of assuming static heterogeneity around dynamic mean models may lead to biased estimates of the population mean, thereby distorting managerial decisions.

Simulated data set.
Consumer Packaged Goods in the Great Recession
We now turn our attention to modeling real choices. Specifically, we model brand choice in the IRI CPG panel data, from January 1, 2006, to December 31, 2011 (Bronnenberg, Kruger, and Mela 2008). We chose this span because it includes the Great Recession, which, according to the National Bureau of Economic Research, began in December 2007 and ended in June 2009 (Business Cycle Dating Committee, National Bureau of Economic Research 2010). Thus, analyzing this time period has the potential to yield purchasing dynamics of interest to both economists and managers. Specifically, we study the evolution of consumers’ individual-level brand preferences, price sensitivities, and feature/display sensitivities across six different product categories: peanut butter, coffee, potato chips, laundry detergent, tissues, and toilet paper. We model the time variation at the monthly level. We retain all panelists who spent at least five times during the data period and save the last four months of data for holdout validation. Summary statistics for the categories are displayed in Table 1.
Summary Statistics for CPG Data, by Category.
Notes: % Ft/Dsp = percentage of observations in which there was at least one brand featured or displayed.
Case study: Preferences for tissues
We focus this analysis on the tissues category and one model: the GPDH logit model with an ARMA mean model. We use this specific example to illustrate the output and insights about dynamic heterogeneity that can be generated from a GPDH specification. The tissues category, in particular, generates interesting patterns of dynamic heterogeneity, and we use the ARMA mean model here as it tended to perform the best among all mean models studied. We defer a discussion of the results across all categories to the next section.
We start with the posterior estimates for the mean model

Mean model for tissues category.
While the mean patterns are certainly interesting, the primary focus of this paper is on capturing how individuals changed relative to those mean trends. In Figure 5, Panel B, we show the difference between the individual-level curves and the estimated mean model,
The nature of the individual-level deviations is determined by the estimated hyperparameters,
We now zero in on a few interesting cases of individual-level evolution that highlight the nuanced insights made possible by considering dynamic heterogeneity. We do this in Figure 6, by focusing on a single parameter, the Brand 2 intercept, which captures the intrinsic preferences for that brand, relative to the baseline, Brand 1. We showcase individuals who spent consistently and whose curves exhibit three interesting patterns: 12

Individual-level dynamic heterogeneity in the tissues category.
Converging: In the leftmost plot in Panel A of Figure 6, we plot a set of individual curves that converge toward the population mean. These customers started in one extreme of the distribution for the Brand 2 intercept, but by the end of the observation window, they were in the middle of the distribution. Under the FO model, these individuals would be estimated as being moderately above or below the population mean, which is true only in the middle of the observation window and does not reflect current or expected future behavior. Crossover: In the center plot in Panel A of Figure 6, we plot a set of customer curves that cross over the population mean. That is, these individuals started out liking/disliking Brand 2 (relative to others) and moved to disliking/liking (respectively) by the end of the observation. Under the FO model, these individuals would be classified as falling near the population mean; in fact, they are perhaps the least average consumers, from a marketing research perspective, as they reflect a strong change in preferences. Diverging: In the rightmost plot in Panel A of Figure 6, we plot individual curves that diverge away from the population mean. These customers started out relatively average in their tastes for Brand 2 but moved to the extremes of the distribution over time. Under the FO model, they would be estimated as being moderately above or below the population mean, which is only true in the middle of the observation window, and again does not reflect current or expected future behavior.
Model fit
We now focus on the results across all six categories to make some generalizations. The key result is that dynamic heterogeneity is pervasive across the six categories. On comparing the FO model to our dynamic heterogeneity model, we find that GPDH fits the data better across all metrics, both in the calibration data and in forecasting tasks, including on metrics that penalize model complexity. We include detailed definitions of these statistics, together with the full set of fit statistics and Bayesian measures such as the Watanabe–Akaike information criterion (WAIC), in the Web Appendix. In Figure 7, we plot a subset of these measures, including in-sample and forecast sensitivity, specificity, and F1 (the harmonic mean of precision and recall), expressed as the lift from using GPDH versus static heterogeneity, across all mean models and categories. The superior fit of GPDH across nearly all of these metrics, both in Figure 7 (lift > 0) and in the Web Appendix, strongly supports our claim that dynamic heterogeneity is present, even in relatively simple panel data sets like grocery store purchases.

Lift from using GPDH.
Parameter estimates and attenuation bias
The hyperparameters of GPDH capture both the magnitude of dynamic heterogeneity for a given parameter, and how much within-individual variation there is, over time. They also allow us to assess the degree by which individual-level trajectories differ from the FO restriction. Across categories, we find that the magnitude of dynamic heterogeneity, η, is typically large, especially for brand intercepts and price sensitivity: for intercept parameters, the mean η is 2 (SD = .64), while for price, the mean η is 2.23 (SD = 2.38). For feature/display, the mean η is .29 (SD = .11).
13
Moreover, GPDH soundly rejects the FO model: the distribution across all categories and coefficients of κ, the inverse length-scale, is centered away from zero, with a mean κ of .03 (SD = .02), and with some values as high as
We found in our simulations that not accounting for dynamic heterogeneity can lead to attenuation bias both in the mean-model estimates and in the overall extent of heterogeneity. We also find empirical evidence of the bias in our real data. Specifically, we find that the empirical standard deviation of individual-level parameters within a given time period is lower when using a FO model than when using GPDH in 75% of cases, with a maximum difference (GPDH SD minus FO SD) of .244 and a minimum difference of only −.034. These results indicate a robust and often substantial downward bias in the spread of FO estimates versus those from GPDH.
Moreover, when we contrast the mean curves recovered from a GPDH specification with those from the FO specification, we see the FO mean curves are biased toward zero. To illustrate this, we develop what we call the signed relative difference (SRD) statistic:
where

Estimated SRDs as a function of the posterior mean estimate of the hyperparameter η.
Individual-level elasticities
Accounting for dynamic heterogeneity is important for accurately computing decision-relevant quantities, including time-varying price elasticities. By both correcting for the attenuation bias and estimating intra-individual dynamics, the individual-level decision variables inferred from GPDH may be dramatically different than those based on a static heterogeneity specification. To illustrate this, we consider own price elasticity of demand across static and dynamic heterogeneity specifications. For each observation in our data, for each brand b, we compute the price elasticity using the standard multinomial logit formula,
First, we consider an illustrative case of a tissues consumer, selected to showcase the differences in elasticities estimated by dynamic versus static heterogeneity. In Figure 9, we present two sets of plots: In Panel A, we show the same consumer’s choice parameters under both dynamic (GPDH) and static (FO) heterogeneity assumptions. In Panel B, we show the implied elasticities over time, for all periods in which the consumer was active.

Illustrative case of a tissues consumer.
Comparing GPDH to FO heterogeneity in Figure 9, Panel A, we see two things: First, the consumer’s brand intercepts deviated substantially from the pattern implied by FO, because of individual-level dynamics. This effect is especially interesting for Brand 2, where the consumer went from negative to positive. Second, we see that the price curve is substantially underestimated using FO, which is likely driven by the attenuation bias. Taken together, these effects produce two effects in the elasticities: First, in almost all cases, the price elasticity is underestimated by roughly 50%. Second, we see the brand intercept dynamics spill over into the price elasticities, with different patterns implied, especially for Brands 1 and 2.
This example demonstrates why we expect to see differences between decision variables under dynamic versus static heterogeneity assumptions. Such differences in elasticities are not limited to special cases. In fact, they are widespread across all categories. To assess these differences more generally, we compute the percentage difference in elasticities from assuming static versus dynamic heterogeneity:
Summary Statistics for the Distribution of
The Great Recession
In the previous sections, we showed the applicability of GPDH to targeting and pricing. In this section, we show how GPDH can also be used by researchers to nonparametrically understand the impact of events, like the Great Recession, on individual-level consumer preferences. In particular, we use our GPDH estimates to understand the changes in individual and market-level preferences during the Great Recession. Researchers have documented how price sensitivity within categories varies with business cycles (Gordon, Avi, and Li 2013) and, more generally, how CPG preferences shifted, on average, during the Great Recession (Cha, Chintagunta, and Dhar 2015). Similar to this previous research, we can use the individual-level GPDH estimates to compute how the average price elasticity of demand changed over time during the recession. We, too, find differences in the effects of the recession on average own price elasticities across the categories, and we include a full discussion of average price elasticities over time in the Web Appendix.
Beyond mean-level analyses, a key benefit of GPDH is that we can also analyze individual-level parameter trajectories. By studying how individuals’ curves deviated from the mean trajectory during the Great Recession, we can nonparametrically analyze how preferences appear to have changed during that period. To illustrate how GPDH individual-level parameter trajectories can be used in this fashion, we created two metrics, related to the timing and impact of the recession: Individual-level maximal rates of change: The first metric aims to understand when preferences changed most rapidly over the course of the data period. To measure this, we again consider the individual-level difference estimates The distribution of the timing of these maximal rates of change then serves as a metric by which we can assess the timing of distributional shifts in preferences. Timing of crossovers: Our second metric isolates the timing of crossovers; that is, the periods in which individual-level curves crossed over the mean curve by either going from the bottom part (half) of the distribution to the top part (half) or vice versa.
15
The distribution of the timings of crossovers then allows us to assess the periods in which preferences appear to have been changing in interesting ways.
Using these two metrics, we find an apparent impact of the recession on individual-level dynamics and the distribution of heterogeneity that differs by category. In Figure 10, Panel A, for instance, we plot the result for the chips category, where we see striking peaks in both metrics associated with the beginning and the end of the recession. Similarly, we find evidence of such peaks in the tissues category. In other categories, most notably coffee, we find no evidence of a recession-era effect, as shown in Figure 10, Panel B. In fact, in the coffee category, as well as in the detergent category, the most rapid changes in the distribution of parameters appears to be concentrated toward the ends of the observation window. 16 While understanding the reasons behind these cross-category effects is beyond the scope of the current work, these findings illustrate the types of analyses enabled by our dynamic heterogeneity framework.

Curve timing results showing the impact of the Great Recession.
Application 2: Dynamic Topic Heterogeneity
Although heterogeneity in marketing has most often been discussed in the context of consumer preferences, GPDH is widely applicable. In this section, we apply it to a different domain: modeling the product-level evolution of review content. In particular, we fuse the latent Dirichlet allocation (LDA) topic model (Blei, Ng, and Jordan 2003) with GPDH to capture dynamic heterogeneity in the evolution of reviews for different products. We apply our model on a data set of time-stamped reviews for tablet computers to address questions such as (1) how the topics used to discuss tablets have changed over time, (2) how the discourse about a focal product is affected by the introduction of new products in the marketplace, and (3) how deviations in product-level topic trajectories reflect the success or failure of the product. Our focus here is on illustrating how our framework can be used across different types of data, contexts, and models, and we therefore do not dwell at great length on the substantive conclusions in this application.
Latent Dirichlet Allocation with GPDH (LDA-GPDH)
Our model extends the standard LDA model of Blei, Ng, and Jordan (2003) to the case where documents pertaining to different groups (e.g., products) evolve over time. In particular, we define a document as the review content of a specific product in a given calendar time period. We model the evolution of the reviews of these products, indexed Generate each topic For each topic
This specification mirrors the calendar time structure used by Dew and Ansari (2018). It captures momentary fluctuations in the prevalence of a given topic, as well as longer-run trends and cyclical variation. We use the squared exponential kernel, which is the limiting case of the Matérn kernel as
with a cycle length, For each product – For each topic – Set the D-th topic unnormalized weight to zero: – Compute the normalized topic assignment probabilities: – For each word token – Draw the actual word token from the assigned topic’s word weights:
In some periods, a given product i may not have any reviews. For that period, the parameters are interpolated or extrapolated.
Comparison with Existing Models
The most common dynamic topic model is that of Blei and Lafferty (2006), which is often referred to simply as the dynamic topic model. In this model the topics evolve over time, but documents are static, and heterogeneity is not accounted for. The focus of this model is solely on modeling the dynamics of content within one group of documents. The LDA-GPDH model is distinct from this classic dynamic topic model in that it focuses on the dynamic evolution of content for multiple groups of documents but assumes that topics are static. The LDA-GPDH framework is thus suitable for the case where new documents are added within each group over time. In the case of reviews, we consider the unit of analysis a single product, where new reviews are continually added over the life span of the product. For other applications, like the modeling of scientific documents within a collection (e.g., theoretical physics papers), considered by Blei and Lafferty (2006), the documents are static. However, the words that are used in documents may change, requiring the evolution of the topics themselves. In other words, LDA-GPDH captures heterogeneity in discourse between groups of documents over time, while the dynamic topic model captures the evolution of content in a single group.
A simpler approach to model the evolution of reviews would be to apply the basic LDA model to documents defined as the composite of all the reviews posted for a given product in a given month. Unlike LDA-GPDH, such an approach treats the reviews of a product as independent across time periods, rather than assuming some consistency of topics within products over time, thus disregarding the primary unit of analysis (the product). As a result, the topics identified by the two approaches are substantially different, with LDA-GPDH finding topics that are consistent within products. For instance, in the case of tablet computers, LDA-GPDH finds many more topics associated with specific brands, while independent LDA finds more topics associated with usage and liking. Moreover, since GPDH shares information across time periods, the topic evolutions estimated using GPDH are much smoother, allowing researchers to better separate noise from true parameter dynamics.
Estimation
As in the choice modeling application, we estimate LDA-GPDH using NUTS. As before, we jointly sample all model parameters, including the individual-level function coefficients, the shared mean function, and the hyperparameters. Unlike in the choice modeling application, LDA-GPDH has discrete parameters, namely the topic assignments, which cannot be sampled by NUTS. Hence, during estimation, we marginalize out the topic assignments, by computing
where
As before, we run the sampler for 400 iterations (200 warmup).
Data
We apply our LDA-GPDH formulation to model the evolution of reviews in a single product category: tablet computers. We use the data from Wang, Mai, and Chiang (2013), which contains the full set of reviews from Amazon for the tablet computer category from September 2003 to July 2012. 17 We limit our sample to the 43-month span from January 2009 to July 2012, which contains the bulk of the reviews (for context, the first Apple iPad was released in April 2010). We further restrict our sample to products that have at least 10 reviews. We aggregate these reviews at the product-month level to form our evolving document stream for each product. For the review content, we follow standard text processing procedures: we first stem the text and eliminate stop words. We then retain all words appearing in at least 5% of observations, but not in more than 75%, where observations are period (month)–product pairs. Finally, we retain the 1,000 words with the highest average term frequency–inverse document frequency scores across documents. This resulted in a data set of 2,686 observations across 265 products.
We ran LDA-GPDH on this data using
Aggregate Results
We start by describing the topics learned by the model, and how their prevalence varies, on average, over time. In Table 3, we show the 10 words with the highest posterior probabilities for each topic. We see that LDA-GPDH identifies meaningful topics that tend to fall into three broad categories: functional topics, capturing aspects of how the products are used or function, especially Topics 1, 3, 7, 10, 12, 13, and 14; experiential topics, capturing consumers’ experiences with their purchases, especially Topic 5; and brand topics, discussing distinct brands and products, especially Topics 2 (XOOM Android), 4 (ASUS Transformer), 6 (Windows), 8 (Samsung Galaxy), 9 (Apple/iOS), 11 (Amazon Kindle), and 15 (HP TouchPad). Note, however, that these distinctions are not always clear: Topic 10, for instance, primarily discusses apps, reading, and downloads but also has discussion of the Kindle; Topics 1 and 9 mention Apple products but also functional words; and Topic 12 has “archo,” which is the stemmed form of Archos, a tablet manufacturer, in addition to multiple functional words.
Summary of the LDA-GPDH model.
Notes: The table shows the GPDH hyperparameter estimates across the 14 unnormalized topics. Higher values of
While the topics are static, the prevalence with which they are discussed changes. Figure 11 plots the mean model

Mean model for selected topics.
which corresponds to the topic weights of the “average” product. We can see in Figure 11 that Topic 5 (Experiential) remained the predominant topic over time, for the average product. Other topics waxed and waned in their prevalence. For instance, we see the relatively recent emergence of Topic 10 (Reading), reflecting the increasing prevalence of this use case in the market. We see the sharp decline in discussion of netbooks and the Windows operating system, reflecting the growing acceptance of tablets as their own product class, with distinct uses from netbooks and personal computers. We also see the rise of the Apple and Samsung topics around the times of their tablet introductions. While these market dynamics make intuitive sense, understanding how individual products evolved relative to these trends is more interesting. We thus turn our attention to understand product-level deviations from these mean trends, which can be recovered from the GPDH specification.
Dynamic Heterogeneity in Topic Weights
As in the choice application, GPDH in our topic model captures individual-level departures from the mean patterns, reflecting in this case product-specific discourse trajectories. The properties of those departures are captured by the two GPDH hyperparameters, which we present for each topic in Table 3. In particular, we find that there is substantially less dynamic heterogeneity for Topics 1 and 5 than for the others. As can be seen in the table, Topic 1 discusses fairly generic tablet-related words, in addition to discussion of the iPhone and iPod. Moreover, as we saw in Figure 11, this topic sharply declined toward the end of the observation window. An interpretation of these patterns is that this topic captures comparisons to iPhones and iPods, which were prevalent before tablets became mainstream; after the popularization of the iPad, these topics were no longer discussed, and thus there was minimal discussion across all brands. Topic 5 captures fairly generic experiential words, and so it is again not surprising that these words occur somewhat more uniformly across brands than other topics. Topics 6 and 12 have the most heterogeneity. Both of these topics reflect somewhat technical language, as well as words associated with niche brands in the tablet space (Archos, Windows). Thus, a large degree of variation in discussion is to be expected.
Market Structure Analysis
The key benefit of using GPDH in topic modeling is that we are able to obtain product-specific topic trajectories: for a given product, how did discourse for that product change, relative to how discourse changed in general? These product-level dynamics can shed light on market structure by examining changes in the discourse for one product during time periods in which potentially competing products were introduced. For example, how did the introduction of Amazon’s Kindle Fire, a highly anticipated Android tablet related to the popular Kindle e-reader, change the discourse in the reviews of existing products? Were niche products affected differently than mainstream products? Was the change in discourse in these products primarily related to brands and products, or did it relate to the functional aspects of the products, too? Answering each of these questions requires understanding how the discourse surrounding individual products changed over time.
In this analysis, we focus on the years 2011–2012, which is toward the end of our observation window. During this time, many next-generation products were introduced, including a new generation of popular Android-based tablets as well as Amazon’s Kindle Fire and two new versions of Apple’s iPad. In particular, this period includes the introduction of Amazon’s Kindle Fire tablet at the end of September 2011 (Period 33 in our data) and the introduction of the short-lived “new iPad” (iPad 3) in April 2012 (Period 40 in our data). 18 To understand how these introductions affected product-level chatter, and what this may imply about market structure, we begin by considering two case studies before reporting results across products.
Case study: Reviews of the iPad 2
Figure 12 shows several of the dynamic topics identified for the iPad 2 (32 GB version). In the figure, we highlight (the leftmost overlaid rectangle) the release window of the Kindle Fire and in green (the rightmost rectangle) the release date of the iPad 3. The first thing of note is the substantial dynamics present during the release of the Kindle. We see that topic weights during that time shifted from Reading to Experiential, as well as to topics about the Apple brand and the iPad. There was also a notable uptick in chatter at that time about Samsung products, reflecting the launch of a new Samsung tablet simultaneous with the introduction of the Kindle Fire, and an uptick in chatter about media playback. After the release of the newer generation iPad 3, we also find changes, this time in chatter about Apple/iOS, about competing Android-based products (e.g., Asus products), and again about reading.

Dynamic topic weights for the iPad 2 (32 GB), estimated by GPDH, for six selected topics.
These patterns reflect different aspects of the tablet market structure. First, while the Kindle Fire was Android-based, it appears to have attracted significant attention even among iPad users, especially when reviews that focus on reading are considered. Previous versions of the Kindle were e-readers but not full-fledged tablets. With the Kindle Fire, Amazon entered the tablet market but retained its emphasis on reading. The ways in which the iPad reviews changed during this period suggest that this move did attract attention, with customers who previously reviewed the iPad for its reading capacity largely vanishing after the Kindle Fire’s introduction. The changes after the introduction of the newest iPad also reveal aspects of the market structure. The uptick in discussion of competing Android products and reading are consistent with a change in the customer base after the new iPad release: customers who continued buying the older version are likely customers who were more price sensitive or who were looking for a tablet with more basic functionality. Thus, we see an increase in chatter about competing but lower-priced brands, as well as a focus on a more basic function (reading).
Case study: Reviews of the ASUS transformer
Figure 13 plots the dynamics over the topics for the Asus Eee Pad Transformer (32 GB version). Immediately, we can see differences between this and the iPad. First, there are clear FO style differences versus the iPad example. For instance, the Asus brand topic is consistently higher for the Asus Transformer’s reviews than it was for the iPad’s reviews. However, there are also clear differences in the dynamics of the topic weights, and departures from the mean-level trends, which are captured by the flexible GPDH framework. For example, similar to the iPad, we also see a rise in discussion of Experiential aspects over the product’s life cycle, reflecting a shift from more functional descriptions of the product to more experiential ones. In terms of responses to product releases, the Asus topics appear to have been substantially affected by the release of the new iPad, but not as much by the release of the Kindle Fire. While there was a noticeable spike in chatter about Reading after the release of the Kindle Fire, it was short-lived. However, after the release of the iPad 3, we see a huge bump in chatter about Apple and the iPad, and a corresponding drop in chatter about the Asus brand.

Dynamic topic weights for the Asus Eee Pad Transformer (32 GB), estimated by GPDH, for selected topics.
Topic dynamics again reflect the tablet market structure. In addition to the dynamics plotted in Figure 13, several other topics are high toward the start of the Asus Transformer’s life cycle, capturing different competing brands or products: Topic 2, about the Motorola XOOM and Android Honeycomb operating system, for instance, started out high. Similarly, Topic 3, about netbooks and touchscreens, started out high, which is especially relevant since the Transformer had an attachable keyboard option. Likewise, Topic 6 (Windows) started out high. Early reviewers emphasized comparisons with these products, reflecting the market position of the Transformer as a mix of these products. The lack of impact of the Kindle Fire’s introduction, together with the seemingly large impact of the iPad’s introduction, also reflects aspects of the product’s use: the Transformer was largely aimed at replacing laptops as a mobile computing system and did not emphasize the reader aspect as much. When the new iPad was released, it likely attracted significant attention from this customer base, as reflected in these changes in topics. Finally, the substantial but short-lived spike in chatter about the Samsung Galaxy also appears to reflect comparisons to that product upon its release, but with no lasting impact, perhaps suggesting different user bases.
Results across products
Finally, we consider patterns of product-level topic evolution across all the products in our data set. Just as in the choice-modeling application, where GPDH allowed us to define new metrics to capture interesting dynamics during the recession, we can also consider new metrics based on dynamic heterogeneity for topic weights, to systematically characterize product-level review dynamics. In particular, we consider a question that was raised in our case studies: during which periods of time did individual products exhibit the most change in topics? To answer this, we consider the following metric:
which reflects the average per period change (from
Conclusion
We developed a novel methodology for capturing dynamic heterogeneity in models of parametric evolution. Across two applications of our GPDH framework to essential marketing tasks, we showed the rich insights that come from modeling the evolution of the distribution of cross-sectional heterogeneity. In our first application, we illustrated the importance of capturing dynamic heterogeneity in choice models with evolving sensitivities, and the managerial and economic insights uncovered by our GPDH specification. In our second application, we showcased the wide applicability of GPDH by employing it in a different context: a topic model capturing the product-level evolution of review content. We applied this model to reviews of tablet computers and used these product-level topic trajectories to shed light on aspects of market structure.
Both applications demonstrate the versatility of the GPDH specification. As GPDH is a way of specifying heterogeneity for the parameters of a focal likelihood, around a particular mean model, it can be used with different likelihoods and dynamic mean models of interest. In this work, we showcased five distinct mean models, ranging from Kalman filter-like state space models to time series specifications and multicomponent GP models. As an extension, we also suggested how to incorporate the drivers of dynamics within mean models, although we did not have appropriate data to demonstrate that use case. In addition, we discussed briefly how endogeneity concerns can be handled via standard control function methods. While we focused only on the benefits of dynamic heterogeneity and thus did not fully explore these extensions, they may be important for researchers interested in adapting our framework in other substantive contexts.
Our work has several limitations that suggest opportunities for future research. First, especially in Application 1, we observe existing customers only. As a result, we cannot rule out the possibility that the observed patterns of heterogeneity are driven by different customer lifetimes (i.e., left censoring). However, the patterns of dynamics we uncover still correctly reflect the changes that occur at a given point in calendar time, regardless of the underlying source of those dynamics. Moreover, given the mature and common nature of the studied CPG categories, we do not expect left censoring to be the primary driver of our results. In addition, because of our emphasis on the methodological contribution of GPDH, we did not fully explore some of the interesting substantive phenomena that were revealed by our GPDH specification. In particular, in Application 1, we noted the prevalence of shifts in brand intercepts versus price coefficients during the Great Recession, as well as the heterogeneous impact of the recession on different categories. In Application 2, we uncovered associations between discourse and product life cycles. Understanding the mechanisms behind these phenomena is beyond the scope of this work but may be interesting topics for future research. Finally, from a computational perspective, our implementation of GPDH using MCMC methods is somewhat slow. Recent advances in Bayesian inference, including variational methods (e.g., Ansari, Li, and Zhang 2018), may prove valuable in accelerating the computation time for these models.
Lastly, while we used GPDH in the context of dynamic heterogeneity, our framework is generally applicable for modeling collections of functions defined on any index, not just time. Other use cases may include spatial modeling and functional modeling of variables. As the modeling of both heterogeneity and dynamics is crucial to marketing, we hope that GPDH will be used and extended for research across a wide variety of domains.
Supplemental Material
Supplemental Material, jmr.17.0056-File003 - Modeling Dynamic Heterogeneity Using Gaussian Processes
Supplemental Material, jmr.17.0056-File003 for Modeling Dynamic Heterogeneity Using Gaussian Processes by Ryan Dew, Asim Ansari and Yang Li in Journal of Marketing Research
Footnotes
Acknowledgments
The lead author gratefully acknowledges the financial support of the American Statistical Association, through its Doctoral Research Award in Marketing; the INFORMS Society for Marketing Science, through its Doctoral Dissertation Award; and the Marketing Science Institute, through its Alden Clayton Award (honorable mention).
Associate Editor
Fred Feinberg
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
