Abstract
This article analyses the upper tails of wealth, income and consumption in India over the period 2012–2018 using rich lists, wealth surveys, income tax returns and consumer expenditure surveys. We find the upper tail to obey a power-law in all three economic resources. Comparing our estimates in 2012—where we possess data on wealth, income and consumption simultaneously—we find that the upper tail of wealth is most concentrated, income slightly less, and consumption is much less concentrated. Unlike wealth and income, the Pareto coefficients for consumption are estimated to have a well-defined mean and variance. Our findings are suggestive of convex saving functions in the income distribution.
Introduction
Although the analysis of economic inequality has become a popularised topic in recent years, it is less apparent how the distribution of wealth, income and consumption are tied together; especially in emerging economies. For instance, is wealth more unequally distributed than income, and if so, then why? Stiglitz (1969) showed that the equilibrium distribution of wealth will be more unequal than the distribution of incomes in the presence of convex saving functions (saving rates go up with incomes); that is, the rich save more, and hence their wealth share is larger than their shares in income. This theory can be tested in data by comparing the parameters of income, consumption and wealth distribution for the rich during the same year. To compare the concentration of resources among (and within) the rich, we need to test whether the distribution function of each resource follows the same probability laws. If the probability distribution is similar, then parametric variation across resources can be used to gauge if the rich save more.
Several economic and social phenomena are known to exhibit power-law type probability distributions (Gabaix, 2016). The distribution of wealth and incomes are canonical examples, because their upper limits are extremely large magnitudes (relative to typical observations). For instance, the Federal Reserve estimates typical US wealth to be around $120k, but there are 22 million individuals worth more than $1 million and at least 600 billionaires (according to Forbes Magazine). Over a century ago, Vilfredo Pareto discovered that the distribution of incomes in Ireland and Britain followed a power-law relationship; the number of individuals with incomes larger than a certain threshold declined at linear rates on a log–log scale. This finding gave birth to the Pareto distribution, expressed as:
In recent decades, several researchers have tried to fit a Pareto distribution to economic data. For example, Dragulescu and Yakovenko (2001) found the upper tails of income and wealth in the United States and Britain to follow Pareto distributions. Klass et al. (2006) showed the presence of power-laws in the Forbes 400 for the United States. Some authors have attempted to confirm the presence of power-laws separately in these economic resources using Indian data for previous years: for example, Sinha (2006) for rich lists in 2002–2004 (100–200 observations), Jayadev (2008) for nationally representative survey data on assets in 1992 and 2002. In this article, we mobilise the largest datasets on Indian wealth, income and consumption to estimate power-laws in the upper tail of the distribution. We use both rich lists and surveys to cover the wealthiest Indians, income tax data to measure top incomes and consumption surveys for the richest consumers. Most importantly, for 2012, we possess data on all three economic resources which allows us to compare the parameters of the Pareto distribution. In both wealth and consumption data, we use more recent innovations in estimation methods to produce estimates of slope coefficients. We confirm power laws in all three economic resources and find that wealth is more concentrated than income and consumption, but the latter has the thinnest upper tail. The estimated Pareto coefficient for consumption is large enough to allow for a well-defined mean and variance. Indian society is very diverse; hundreds of billionaires coexist with the world’s largest number of poor people; thus, the presence of power-laws in the upper tail is indicative of statistical regularities that transcend the level of economic development. Given that the concentration of wealth is larger than income, and income is more concentrated than consumption, the data favour convex saving functions in Indian incomes as a possible source of rising inequality.
Wealth Distribution
Rich-list Data
Forbes Magazine produces annual lists of the richest individuals in several countries. Starting in 2012, Forbes began compiling rankings of the richest 100 individuals in India. In the past, these rankings were small, between 1 and 46 individuals. We scraped these lists from the website of Forbes Magazine for 2012 and 2018. As we will show in a moment, these 2 years of data also allow us to compare and pool with more nationally representative survey data. The data present us with the names of India’s wealthiest individuals and their respective wealth in US dollars. We converted these monetary figures into Indian rupees (₹) and adjusted for inflation using the consumer price index (CPI).
In the presence of a power-law, the relationship between rank (

Rank Ordered Plots of the Richest 100 Indians for 2012 and 2018 on a Double Log Scale.
Survey Data
Next, we estimate the same rank-size of wealth relationship in nationally representative wealth survey data for 2012 and 2018. India’s official statistical agency publishes unit-level data reporting assets and liabilities in the All-India Debt and Investment Survey (AIDIS). From these surveys, we computed wealth2 by subtracting financial liabilities from the sum of assets (all in monetary amounts), as per convention. Compared to rich lists, where the main interest is the opulence among the wealthy, the AIDIS is meant to analyse debt and credit situations across different socio-economic groups. The data are top-coded, and most wealth surveys across the world tend to underestimate the upper tail due to sampling error, under-reporting and non-response correlated with wealth (Davies & Shorrocks, 2000; Vermeulen, 2018). The estimates that follow should be understood with these limitations.
We estimated wealth per adult and converted all monetary aggregates to 2012 prices using CPI. Because these data are nationally representative, the entire distribution is unlikely to follow a power-law. Clauset et al. (2009) developed a method to estimate a cutoff
In Figure 2, we show plots for both years using the truncated sample from the wealth survey. Based on this sample, we found excellent fits to the power-law, with R-squared of 0.988 and 0.99 in 2012 and 2018, respectively. Resonating with our previous findings, we note that the Pareto coefficient (a) increases over time from 1.66 (2012) to 1.86 (2018). Therefore, wealth concentration within the survey upper tail decreased over 2012–2018. We observe, however, that the extreme upper tail deviates from the power-law because of top coding. The highest ranks in 2012 are flatter than indicated by the fitted line, and steeper in 2018.

Rank Ordered Plots of the Upper Tail in Survey Data for 2012 and 2018 on a Double Log Scale.
Wealth Pooled Across Both Data Sources
Both previous datasets have shown the upper tail of Indian wealth in 2012 and 2018. However, the sample quite clearly underestimates the magnitude of wealth. To be sure, the measurement of wealth is much more challenging than income and monetary value of consumption expenditures. Asset values are sensitive to valuation, and there tends to be variation in the thickness of their respective markets, based on geography or asset class. Regardless, the gap between the wealthiest observation in the surveys and the 100th ranked individual on the Forbes list is very large—around two orders of magnitude. To estimate a more plausible upper tail, we pool the rich list and the upper tail of the survey. To account for sampling weights, we allocated a weight of 1 to all individuals3 on the rich list. If the cell mentions more than one individual (e.g., Person A and family), then we allocate a weight of 2. Observations from the survey were left unchanged, and rank adjustment was applied to estimate a precise slope.
Our results from pooled wealth data are shown in Figure 3. We note interesting developments in comparison to estimates from separate datasets. For the pooled upper tail of wealth in India, we find the Pareto coefficient declines over time, from 1.29 in 2012 to 1.21 in 2018. We thus conclude that wealth concentration increased in the upper tail as a more plausible scenario (i.e., including the rich from the top-coded survey and rich lists). The data also fit the power-law with R-squared of 0.987 and 0.969 in 2012 and 2018, respectively, thus improving the fit provided by rich-list data alone. All our results are summarised with the appropriate statistical indicators in Table 1.

Rank Ordered Plots of the Upper Tail in Pooled Survey and Rich-list Data for 2012 and 2018 on a Double Log Scale.
Pareto Coefficients Estimations for Wealth Using Rank-size Regression on Survey and Forbes Data.
We next follow the same steps to estimate the presence of power-laws in consumption surveys. We used the latest (2012) consumption survey published by the National Statistical Organisation. These surveys are conducted regularly, although recently there was some controversy about retraction of the 2017–2018 survey in the Indian media. The data are unit-level observations reporting household consumption levels using a two-stage stratified sampling strategy. We converted all expenditures from monthly to annual frequency to allow comparison across datasets. As we mentioned previously for the official wealth surveys, consumption surveys also suffer from top coding, sampling errors and under-reporting or misreporting issues. Often, households forget to mention some expenditure, thus introducing noise across observations. We proceed with these data limitations in mind. We applied the Clauset et al. (2009) procedure to identify the power-law subsample within the data. This truncated the sample to 4,188 observations to which we applied
Pareto Coefficients Estimated for Consumption Using Rank-size Regression on Survey Data.
Pareto Coefficients Estimated for Consumption Using Rank-size Regression on Survey Data.

Rank Ordered Plots of the Upper Tail in Consumption Data for 2012 on a Double Log Scale.
Last, we estimate the upper tail income distribution data for India. Our income dataset is different from the surveys used in previous sections. We obtained income tax returns published by the official tax authorities in India. These reports have been published for every year from 2012 to 2018 on the website of the income tax department of India and are easily translated into excel sheets using online software. Due to the legal requirement to file taxes and payroll deduction at source, these tax data provide excellent coverage of the upper tail. For every year, the income tax department’s tax tables report the number of tax filers as a function of income brackets. For instance, in 2012, there were 7.69 million tax filers who reported income between ₹150k to ₹200k, 4.5 million tax filers between ₹250k and ₹350k and so on. However, we are only interested in the distribution of income among individuals and families (the website also reports tax filings of companies). For each tax filer, income totals include the sum of wages and salaries, capital incomes, capital gains, rents and business incomes. We cumulated the number of tax returns above the minimum of each progressively increasing bracket and divided cumulated returns by the total adult population. Thus, our tables report
Table 3 reports our estimates of the Pareto coefficient in each year. First, we note that a very small fraction of the Indian population makes sufficient income to file taxes. In 2012, only the top 2–3% of the adult population were liable to file taxes. Due to inflation (bracket creep) as incomes increase—but tax brackets remain the same—more people become eligible to file taxes. Additionally, the digital capabilities of the tax authorities have increased over time. Based on these factors, the proportion of tax filers in the adult population increased to around 8% by 2018. Regardless, we found excellent fits to the Pareto distribution between 2012 and 2018 with R-squared approximately 0.99 in all years. The Pareto coefficient was estimated to be between 1.59 in 2012 and 1.66 in 2018, having been slightly larger (1.7) in 2015 and 2017. Because the fluctuation in
Pareto Coefficients Estimated for Consumption Using Rank-size Regression on Survey Data.
Pareto Coefficients Estimated for Consumption Using Rank-size Regression on Survey Data.

Complementary CDF Plots of Income-tax Data, Pooled for 2012–2018 and Adjusted for Inflation Using CPI (2012).
The results in this article suggest that the upper tails of wealth, income and consumption obey a power-law probability distribution using data from India over the period 2012–2018. In each dataset, we found good fits to the Pareto distribution, either using the rank-size specification
The relatively more equal distribution of consumption can be understood as resulting from convex saving functions:
Finally, the fact that wealth is most concentrated in India is not surprising. While incomes may be a function of education and effort, wealth can be transmitted across generations and rates of return may increase systematically with wealth. Second, econophysicists propose that the evolution of wealth among the rich follows a multiplicative process, rather than additive processes which are more typical of the incomes of the middle class (Castaldi & Milaković, 2007; Sinha, 2006; Tao et al., 2019; Yakovenko & Silva, 2005). Another important factor is the prevalence of primogeniture where the firstborn son inherits all wealth (Stiglitz, 1969), or even a contentious split of inheritance. In a recent study of India’s corporate sector (Banaji, 2022), it was found that several of India’s largest (and longstanding) fortunes disintegrated after they were split among heirs. The distribution of income can also become Pareto-distributed due to superstar effects in labour incomes (Gabaix, 2009)—the salaries of CEOs, film stars and athletes, for example—and capital incomes following similar processes to wealth (Shaikh, 2017).
Footnotes
Acknowledgements
The author thanks Victor Yakovenko for insightful comments, suggestions and feedback that helped improve this manuscript. All remaining errors are my own.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author received no financial support for the research, authorship and/or publication of this article.
