Abstract
The cost of living varies as much across locations as it does over time. We demonstrate the importance of considering locational cost of living differences in empirical models of the demand for state lotteries. Previous research has shown that the nominal-income elasticity of demand for lottery tickets is less than one, suggesting that individuals and geographic regions with lower incomes tend to have a greater percentage of their income allocated toward lottery ticket purchases than do wealthier individuals and geographic regions. We first provide a conceptual framework that reveals that real-income elasticities generally will be different from nominal-income elasticities. We then reestimate traditional cross-sectional models of lottery demand using a sample of metropolitan statistical areas. We find that the magnitude of income elasticity estimates is smaller when local cost of living is omitted from empirical models, especially in the case of instant lottery games.
Research has shown that individuals and geographic regions with lower incomes tend to have a greater percentage of their income allocated toward lottery ticket purchases than do wealthier individuals and geographic regions (see, e.g., Clotfelter and Cook 1987, 1989; Scott and Garen 1994; Farrell, Morgenroth, and Walker 1999; Price and Novak 1999; Forrest, Gulley, and Simmons 2000; Garrett and Coughlin 2009). As a result, the distributional burden of expenditures on lottery tickets is generally characterized as regressive. 1 The distributional burden of lottery ticket expenditures is traditionally determined by estimating the income elasticity of demand for lottery tickets, with a value less (greater) than one, indicating regressivity (progressivity). To determine the income elasticity of demand for state lotteries, researchers typically estimate an equation where lottery ticket sales is a function of income and other demographic characteristics. Many studies use cross-sectional data, and the unit of observation is a zip code, a city, a county, or a state because individual-level data are not usually available. As a result, the “change in income” comes from the variation in nominal income across locations with no regard for price-level differences across locations.
We argue here that the empirical models of lottery demand described earlier should include some measure of the local price level in order to capture differences in purchasing power across locations. It is standard procedure to adjust monetary variables (such as income, revenue, and gross domestic product [GDP]) for inflation when comparing them over time. The point of such an adjustment is to ensure that differences in monetary variables over time reflect only differences in purchasing power. Yet, it is common practice to compare only nominal incomes in comparisons of locations across space at a given time despite the fact that price-levels, and thus the cost of living, across locations are as varied as they are across time. For example, the Consumer Price Index (CPI) in 1985 (107.6) and 2000 (172.2) yields the same cost of living difference as the cost of living difference between Amarillo, Texas (91.6), and Los Angeles, California (146.8), in 2000. It should be of little surprise that the cost of living varies widely across locations. According to the data from the 2000 census, the median price of a house in San Francisco was five times greater than the median price of a house in Pittsburgh. Significant cross-city variation in housing prices exists even after adjusting for the quality of housing (Gabriel and Rosenthal 2004; Chen and Rosenthal 2008). In addition, various cost-of-living indexes (COLIs) indicate that prices of other consumption goods (groceries, utilities, transportation, and health services) vary across locations. 2
Several recent studies have demonstrated the importance of considering local cost of living. Black, Kolesnikova, and Taylor (2009, 2014) show that differences in local prices—namely, housing prices—help to explain significant differences in the college wage premium across cities and the evolution of income inequality in the United States. Moretti (2011) emphasizes the importance of distinguishing between nominal and real earnings and shows that at least 22 percent of the increase in the college wage premium over the past thirty years is explained by spatial differences in the cost of living. Albouy (2009) investigates the unequal burden of federal taxation across cities that results from the differences in locational cost of living and wages. The cost of living has also been found to be an important determinant in the demand for children (Black et al. 2013).
Our objective in this article is to estimate empirical cross-sectional models of lottery demand that have been established in the literature (Mikesell 1989; Hansen 1995; Price and Novak 1999, 2000; Garrett and Marsh 2002; Ghent and Grant 2010), but contribute to these models by accounting for cost of living differences across locations (the cross-sectional units of observation) and examine how estimates of the income elasticity of demand change when cost of living differences are considered. We first demonstrate conceptually that the failure to consider differences in the cost of living when estimating lottery demand may yield an income elasticity of demand that is quite different from that obtained when the cost of living is considered. As we show later using a sample of Metropolitan Statistical Areas (MSAs), our income elasticities of demand based on nominal incomes and nominal lottery sales are generally lower than the income elasticities that consider cost of living differences across MSAs. To see the importance of considering cost of living differences, assume as an extreme example that the demanded quantity of lottery tickets is the same in two cities but income and the cost of living in city 1 are twice as high as those in city 2. If we ignore the cost of living when calculating the income elasticity of demand, we would conclude that lottery demand in this society is absolutely income inelastic. This is, however, an erroneous conclusion, as we really cannot say anything about the income elasticity of lottery demand by observing these two cities since there is no variation in purchasing power between the two cities. We further motivate this observation in the next section of the article that discusses our conceptual framework.
Conceptual Framework
One of the main obstacles facing researchers in estimating the income elasticity of demand for lottery tickets is the lack of individual-level data. As a result, the estimation must rely on aggregate-level data. However, it is important to realize that an “income elasticity of demand” obtained from aggregate data is conceptually different from the textbook definition of income elasticity of demand that is based on individual preferences and utility maximization. Thus, an aggregate measure of income elasticity cannot be assumed to represent income elasticity of demand at the individual level. Despite this arguably misleading terminology, the previous studies (and our study as well) find that an aggregate income elasticity estimate remains useful for examining the response of regional sales with respect to changes in regional income as long as interpretation at the individual level is not made.
With cross-sectional data in hand, the income elasticity of demand for lottery tickets (β1) has traditionally been obtained by estimating the following equation (see Mikesell 1989; Hansen 1995; Price and Novak 1999, 2000; Garrett and Marsh 2002; Ghent and Grant 2010):
where i typically represents a zip code, a county, or a state, depending on the study; Xi
is per capita lottery sales; Yi
is per capita income; and the matrix
Of course, because the cost of living varies significantly across locations, the same nominal income does not imply the same purchasing power across locations. One way to address the issue is to consider real income, as is common practice in making comparisons over time. We argue that similar adjustments must be made when comparing incomes across space. In this case, we calculate real income as the ratio of nominal income and a measure of the local cost of living, following Moretti (2011). A similar adjustment is made for lottery sales. Equation (2) then can be used to estimate the real-income elasticity of demand for lottery tickets:
where COL i is the local cost of living in location i, Yi /COL i is real income for location i, and Xi /COL i is real lottery sales for location i. The estimated coefficient β2 is real-income elasticity of lottery demand.
It is important to understand the difference between equations (1) and (2). In equation (2), income now reflects the purchasing power of income across locations, and lottery sales now are presented in comparable dollars across locations. Note that this adjustment is analogous to the time-series case—just as we would not compare nominal incomes in 1950 and 2000 without adjusting for changes in the price level over time, we should not compare nominal incomes in, say, New York City and Omaha, Nebraska, without adjusting for differences in the cost of living. We want to have comparable measures of the purchasing power of income across locations, just as in time-series analysis where we want to have comparable measures of the purchasing power of income across time.
Similarly, comparing sales of a good across locations also requires adjusting for the local price level, as what matters is the cost of the good relative to all other goods. Another way to think of this is that the purchasing power of the sales revenues from a good (which is equal to consumer expenditures on the good) is different between two locations. Again think of New York and Omaha: US$10,000 of sales revenue in New York has less purchasing power than US$10,000 of sales revenue in Omaha.
Because purchasing power is what matters, a cross-sectional regression should therefore adjust monetary variables for price differences across locations to ensure that the monetary variables are comparable across all observations in the sample. If no adjustment for price differences across locations is made, the interpretation of the estimated effect of one monetary variable (e.g., income) on another (e.g., sales) becomes less clear as the estimated effect (coefficient) is based on units of observation that are not directly comparable across locations. This is analogous to the time-series case where nominal units of observations are not directly comparable across time due to temporal price-level differences.
Although the previous discussion highlighted the importance of controlling for cost of living differences across locations, what remains is to think conceptually about why estimates of the nominal elasticity (β1) and the real-income elasticity (β2) may be different. The following exercise provides insight into the potential difference in the nominal-income elasticity of demand estimated from equation (1) and the real-income elasticity of demand estimated from equation (2). For simplicity, let there be only two locations L and H—location L has low nominal income

The relationship between demand and income—real versus nominal income.
Now consider that the cost of living differs in the two locations. What happens to the income elasticity when we present income and sales in real terms? Real income for each location can be computed as nominal income divided by cost of living, so
The same conversion to real dollars can be done for sales of good X. So we have
The slope of each line represents one of the two possible values for the real-income elasticity (ηR1 and ηR2) for good X, but ηR1 is greater than ηN while ηR2 is less than ηN. To see this, first compare ηR1 and ηN. From the figure, it is clear that
The main point from figure 1 is that the real-income elasticity of demand can be larger or smaller than the nominal-income elasticity of demand depending upon the relative differences in the percentage change for real income and real sales. That is, adjusting for the cost of living alters the percentage changes between the two sales and the two income observations.
It is not difficult to show that the nominal-income elasticity of demand is equal to the real elasticity of demand if and only if the ratio of expenditure on a good to income is constant across locations. 8 In other words, we can ignore differences in the cost of living when estimating the income elasticity of demand only when the income share spent on a good (lottery tickets, in our case) is the same in every location. 9 Since it is unlikely that this condition holds, the estimated nominal- and real-income elasticities will be different. 10 Just how different the elasticities are is an empirical question, which we explore in the remainder of this article.
Data
To estimate the distributional burden of state lotteries, one would ideally prefer data at the individual level. However, no nationally representative sample of individuals exists for the United States that provides information on individuals’ lottery expenditures, geographic location, incomes, and demographic characteristics. 11 Thus, we follow the majority of the existing literature on lottery demand and use cross-sectional data at the most disaggregated unit of observation available—in our case, MSAs.
We would like to have an index similar to the one provided by the Bureau of Labor Statistics (BLS)—the Consumer Price Index for All Urban Consumers (CPI-U). The CPI-U represents a weighted price index of a basket of goods and services. Unfortunately, the BLS produces local indexes for only twenty-seven MSAs. More importantly, even these local indexes are not suitable for comparing living costs across areas as they measure only how much prices have changed over time in a given MSA.
Instead, we use the American Chamber of Commerce Research Association (ACCRA) Cost of Living Index (COLI) provided by the Council for Community and Economic Research. 12 The COLI measures relative price levels for consumer goods and services and has been used in previous studies to capture price variation across locations (Cebula and Coombs 2008; Nonnemaker et al. 2009; Moretti 2011). The price index for each MSA is interpreted as a percentage of the average for all urban areas (where the average is set to 100). 13 The COLI is a weighted average of price indexes of six different categories—namely, groceries, housing, utilities, transportation, health care, and miscellaneous good and services. 14 Each category includes many goods, but it should be noted that the basket for COLI is smaller than the one used by the BLS for computing CPI-U. The big advantage of COLI is that it allows comparison between MSAs at a given point of time and is available for a larger number (over 300) of MSAs. Thus, we use an MSA as a unit of analysis because it provides the lowest level of aggregation for which local price data are available.
Lottery-ticket sales at the MSA level are not readily available and therefore had to be constructed from county-level lottery sales data. To do so, we first obtained a list of all counties within each MSA (for which we had obtained local price data) by using the 2000 census MSA boundary definitions and component (county) names. 15 We then contacted state lottery agencies and obtained county-level sales data for instant lottery games (scratch-offs) and online lottery games (e.g., Lotto, Mega Millions, and Powerball) for the year 2000, and then summed sales at the county level to arrive at online lottery sales at the MSA level and instant lottery sales at the MSA level. Examining online games and instant games separately allows us to explore the role of local prices in explaining the demand for different lottery products, as well as providing us with additional tests of our hypothesis.
Online games and instant games are considered different lottery products because online lottery games offer much higher jackpots than instant lottery games, and the potential frequency of play for online games is less than that of instant games as drawings for online games are aired on television only several times a week. Research has shown that the income elasticities of demand for online games and instant games can be different. 16
Descriptive statistics for COLI and lottery sales are shown in table 1. Perhaps not surprisingly, there is significant variation in COLI across MSAs. The Bryan/College Station MSA in Texas has the lowest COLI of 88.9, and the New York MSA has the highest index of 242.7. There is also large variation in lottery sales across MSAs. Per capita, nominal (real) instant sales range from about US$21 (US$18) to US$131 (US$140); per capita, nominal (real) online sales range from US$13 (US$14) to US$178 (US$150). This sample variation is similar to what typically exists in cross-sectional samples of lottery sales for states, counties, and zip codes. We include state dummy variables in our empirical models to capture some of the variation in lottery sales across MSAs.
Descriptive Statistics.
Note: Sample size is 111 Metropolitan Statistical Areas (MSAs) for the year 2000. See Appendix for a list of MSAs. COLI = Cost-of-living index. aThe McAllen–Edinburg–Mission, TX, MSA had the lowest per capita income of all 276 MSAs in the fifty states for the year 2000.
We follow the past literature and also include several economic, demographic, and game characteristic variables in our models of lottery demand. The economic and demographic variables we include are per capita personal income, population density, and the percentage of the population with a bachelor’s degree or higher. 17 Game characteristics include the age of the lottery in years, an indicator dummy variable for whether the state participates in multistate lottery games, the number of years the state has participated in multistate lottery games, and an indicator variable for whether the state has commercial casino gambling. 18 The values for these variables are the same for each MSA in a state since the lottery is statewide. The age of the lottery captures the differences in each state lottery’s life cycle (Mikesell 1994). Because we are comparing different state lotteries in different stages of their life cycles, the expected sign on age is ambiguous. Multistate games (e.g., Powerball) generate the largest jackpots, and thus states that participate in these games are expected to have higher sales (Garrett and Sobel 1999; Kearney 2005). The casino dummy variable captures any effects of competition between casino gaming in the state and the state lottery (Elliot and Navin 2002; Garrett and Coughlin 2009). Descriptive statistics for these variables are shown in table 1.
We conduct our analysis using data on lottery sales, local prices, personal income, and demographic and game characteristics for 111 MSAs for the year 2000. The sample size and year of study were dictated by the greatest availability of local price and lottery sales data. 19 The MSAs used in the analysis are listed in Appendix A.
Empirical Results
We estimate equations (1) and (2) using per capita instant lottery sales and per capita online lottery sales as our dependent variables. All equations contain the aforementioned economic, demographic, and game characteristic variables, as well as a set of state dummy variables to capture potential heterogeneity across states. 20 Because we wish to obtain the income elasticity of demand, lottery sales and per capita income are converted to natural logarithms before estimation.
For each game category, we compare the nominal-income elasticity estimates from equation (1) with the real-income elasticity estimates from equation (2). As our previous discussion suggests, we expect to find that the income elasticity coefficient from equation (1) is different than the income elasticity coefficient from equation (2). This would support our hypothesis that local prices play an important role in explaining cross-sectional differences in lottery demand and that the failure to include local cost of living in models of lottery demand can result in income elasticity estimates that are biased, thus providing an inaccurate estimate of the distributional burden of lottery ticket expenditures.
The empirical results from our models of lottery demand for instant sales and online sales are shown in tables 2 and 3, respectively. All equations were estimated by Generalized Least Squares (GLS) using White’s heteroscedasticity-corrected standard errors. 21 The results presented in columns (1) and (2) of tables 2 and 3 are from equation (1), which estimates the nominal-income elasticity of demand excluding control variables (column 1) and including control variables (column 2). The key feature of these “traditional” models is that they omit local cost of living, thus failing to capture how differences in the purchasing power of income influence lottery sales across MSAs.
Instant Lottery Sales.
Note: White’s heteroscedasticity-consistent standard errors are shown in parentheses. The coefficient on population density is multiplied by 1,000. Number of observations = 111. Unit of observation is Metropolitan Statistical Area (MSA).
*Denotes significance at 10 percent. **5 percent or better.
Online Lottery Sales.
Note: White’s heteroscedasticity-consistent standard errors are shown in parentheses. The coefficient on population density is multiplied by 1,000. Number of observations = 111. Unit of observation is Metropolitan Statistical Area (MSA).
*Denotes significance at 10 percent. **5 percent or better.
Consider the regression results from equation (1) with control variables, which are shown in column (2) of each of the two tables. For instant lottery sales (table 2), we estimate a nominal-income elasticity of 0.268 that is not statistically different than 0. The estimated income elasticity for online games (table 3) is 1.106 and is statistically different than 0, but it is not statistically different than 1. This finding is similar to that of Mikesell (1989) who used county-level data for Illinois and Perez and Humphreys (2011) who used individual-level data for Spain.
Next, we estimate equation (2) where the dependent variable is real sales per capita and the key independent variable is real per capita income. The regression results are presented in columns (3) and (4) of tables 2 and 3 where control variables are both excluded (column 3) and included (column 4). As expected based on our earlier conceptual framework, the real-income elasticity of demand is different than the nominal-income elasticity of demand in each case. 22 Specifically, the estimated income elasticity is larger when we consider real income than when we consider nominal income, especially for instant lottery sales.
Consider specific regression results. For instant lottery sales, the nominal-income elasticity is 0.268 (column 2 of table 2) and is not statistically different than 0, whereas the real-income elasticity for instant lottery sales (column 4 of table 2) is 0.817 and is not statistically different than 1. For online lottery sales, the nominal-income elasticity estimate (column 2 of table 3) is 1.106 and the real-income elasticity (column 4 of table 3) is 1.260, both of which are statistically different than 0 and not statistically different than 1.
In agreement with our predictions, the results reveal the nominal-income elasticity of demand that omits local cost of living can be quite different from the real-income elasticity, which considers the local cost of living. We find that for instant sales and online sales, the real-income elasticity is greater than the nominal-income elasticity in each case. The difference in the nominal and real elasticities is more pronounced for instant lottery sales. The local cost of living appears to play an important role in determining the income elasticity of demand for lottery tickets.
Concluding Comments
A growing body of literature argues that empirical modeling of cross-sectional data should account for geographic variation in the cost of living, much in the same way that time-series data are frequently adjusted for inflation in order to make accurate comparisons over time. Accounting for differences in purchasing power across locations when using cross-sectional data is just as important as accounting for differences in purchasing power across time.
In this article, we explored the role of geographic variation in cost of living in empirical models of demand for state-lottery tickets across MSAs in the United States. Previous research on the distributional burden of lottery ticket expenditures has used only nominal income and nominal lottery sales across locations. We argued that the failure to consider locational cost of living differences in empirical models of lottery demand may yield incorrect estimates of the income elasticity of demand. Specifically, our conceptual framework suggested that an estimated nominal-income elasticity of demand will be different than the real-income elasticity of demand, except in the very special case when the share of income devoted to lottery expenditures is the same in all locations.
In accordance with our conceptual framework, our estimated income elasticities are different when controlling for cost of living differences across MSAs. Our estimated real-income elasticities are larger than nominal-income elasticities, and thus reveal that the regressivity of state lotteries may be overstated when locational differences in the cost of living are ignored. Our results suggest that if individual-level data on lottery expenditures were available, then cost of living differences across individuals’ locations should be considered in these individual-level models of lottery demand as well. 23
Our analysis of lottery demand across MSAs demonstrates the importance of accounting for differences in the cost of living (or prices, more generally) across geographic locations. Although we focused on the single example of the demand for lottery tickets across MSAs, the conceptual framework and its conclusions presented here are applicable to any cross-sectional analysis using monetary variables (e.g., government spending, sales, wages, etc.) at any level of geographic aggregation (e.g., zip codes, counties, cities, and states). 24 Since we find that controlling for geographic variation in the cost of living matters for lottery demand, it is reasonable to believe that doing so also matters for various other issues as well, such as tax incidence, impact of intergovernmental grants on local spending, gender and racial wage gaps, income and tax inequality, and union wage premiums, to name just a few. Given that research in these areas often has important policy implications, it may be beneficial for future research to revisit previous analyses and consider that accounting for the geographic variation in the cost of living may provide alternative results and policy recommendations.
Footnotes
Appendix A
The following Metropolitan Statistical Areas (MSAs) are used in the analysis (2000 US Census definitions).
Appendix B
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
