Abstract
This article provides a statistical decision-making framework for the Gini index through the introduction of practically usable confidence intervals for this index. The resultant intervals enable hypothesis testing with the Gini index as well as providing a mechanism for studying the reliability of its estimates. The article presents these confidence intervals and demonstrates their prospective uses with an illustrative example to officially published estimates for the index from 24 countries.
JEL Classification: C02, C13
Keywords
INTRODUCTION
On reviewing the estimation practices of the Gini index, Karagiannis and Kovacevic (2000: 119) observed that it is common practice for its margin of error to be rarely if ever reported for the main reason that the margin’s computation is both computationally challenging as well as expensive to do since no readily available technique exists for its derivation. Ogwang (2000: 123) likewise found the same, observing that it is not yet commonplace for the index’s margin of error to be reported. As Ogwang (2000: 123) explains:
This is because most estimators for the … errors of the Gini index that have been suggested so far … are either mathematically very complicated or require heavy computation which cannot be conveniently undertaken using commonly available … software packages.
This premise appears to be somewhat resilient, as another review by Giles (2004: 425) suggests. In explaining the non-reporting of the margin error for the Gini index, Giles (2004: 425) notes that:
This is because most of the formulations … that have been proposed are mathematically complex, or they are computationally intensive.
Considering that a statistic’s confidence interval is defined by the range of its estimate, as determined by the deviations from its margin of error, the implication of the foregoing situation is that effectively the Gini index is typically estimated without confidence intervals. The aim of the present article is to propose an improvement to this situation, by relying on an enquiry from Mandelbrot (1997: 216–7), which suggests that the problem is not as constraining or insolvable as it appears. The trick lies in realising that the familiar McKay approximation for the sample coefficient of variation (c) actually extends to the sample Gini index (G). This is due to two established results of mathematical statistics: the De Vergottini inequality for the coefficient of variation and the Glasser inequality for the Gini index. Using these inequalities, in the present article McKay’s approximation is extended to the Gini index to yield the extended McKay confidence intervals for this index. This will be done by asymptotics—being the conventional mathematical approach to extending a statistical technique from one estimator to another (Tukey, 1986: 74). Under this approach, an extension of a statistical technique from one to another estimator is achieved when an asymptotic equality between the estimators is identified to exist. To remind, an asymptotic equality is an approximate equality whose relative error—in terms of the percentage difference among its participating estimators—becomes infinitely small as the number of observations increases, when in the limit with an infinite number of observations this error disappears, thereby leading to a strict equality.
In the present case, the demonstration of asymptotic equality will focus on showing that the coefficient of variation is asymptotically equal to the Gini index. Resultantly, the confidence intervals for the sample coefficient of variation from McKay’s approximation are equally applicable to the sample Gini index, simply by replacing the coefficient of variation with the Gini index in their formulation. As part of the intended demonstration, the Section 2 of this article deals with the De Vergottini and Glasser inequalities in sequential order, as knowledge of the former is needed to understand the latter. Thereafter the extension to McKay’s approximation is given in Section 3, and the practical applications of this are considered in Section 4. In the end, a conclusion to the presented work is offered in Section 5.
In order to keep mathematical demonstrations simple, a minimal amount of basic algebra is deployed in order to simplify many procedures, which should at the same time make them more meaningful. As part of this, the sample size (n) is taken to be equal to the population membership (N) or n = N. As Glasser (1962a: 628) explains, first, the meaning of this assumption is that every population member is sampled, such that the possibility of mathematically dealing with incomplete information is excluded and second, that samples of any size are being considered, as the population number varies.
Before delving into the detail above, it should briefly be noted that the present situation is different from that described by Bertrand et al. (2004: 254), of statistics whose margins of error are on the first place computable and thus available, even though they might be biased. As Bertrand et al. (2004: 254) noted, the issue in such cases becomes how to reduce or remove this bias from a statistic’s estimate. The issue in the present situation is to actually find an accessible computational method for the error margin to begin with, through the derivation of a confidence interval. Furthermore, provided a statistic has confidence intervals, then the determination of bias and its elimination from its estimates is also no longer a matter of concern. This is because half the width of a confidence interval gives an estimate of the bias in the estimation of a statistic. In turn the range of possible bias-free estimates for a statistic is obtained by correcting the statistic’s estimate with the product of its estimated bias. An illustration of this to the present instance will be given in Section 4 of this article.
THE DE VERGOTTINI AND GLASSER INEQUALITIES
The De Vergottini inequality is named after De Vergottini (1950: 452), who introduced it as a fait accompli result with the interesting property that it leads to an asymptotic equality between the coefficient of variation and the Gini index, which has an exact solution of one-third. Most recently, Piesch (2005: 284) reproduced the De Vergottini inequality as a well-known result without deriving it. By contrast, as part of building-up the present argument, a simple derivation of this inequality is offered, being based on its solution in the limit, as the number of observations increases.
To begin with, we know that the coefficient of variation for the ranks of the data (i) is the ratio between their standard deviation (vi) and mean (ni). If we square this ratio, we get:
Taking the square root of expression (1) results in the following re-expression:
As the number of observations increases, from 25 onwards, the square root term involving the number of observations in expression (2) approaches 1, leaving an approximate solution for the coefficient of variation for the ranks of the data as the inverse of the square root of 3. The same result also follows if we derive the coefficient of variation for the values of the data by a regression through the origin between the data’s values and ranks. In the context of a regression, the standard deviation of the data’s values (vx), which represent the dependant variable, is the product between the slope of the regression (b1) and the ratio between the standard deviation of the data’s ranks (vi), which represent the independent variable, and the correlation between the values and the ranks (t). In short:
The correlation between values and ranks in expression (3) is readily recognisable. It is the familiar Stuart correlation, whose maximum value as per the Stuart inequality is 1 or t ≤ 1 & t = 1. In addition from expression (1) we already know what the standard deviation for the ranks of the data is, namely,
Since we are working with a regression through the origin where the intercept is zero, the mean of the values is given by:
The mean of the ranks in expression (5) follows from expression (1). Taking the ratio of expressions (4) and (5), followed by cancelation and the collection of power terms, leads to the coefficient of variation of the data’s values:
By increasing the number of observations from 25 onwards, the square root term involving the number of observations in expression (6) approaches 1, this time leading to an approximate solution for the coefficient of variation for the values of the data as the inverse of the square root of 3. Clearly the ratio between expressions (2) and (6) is 1.
But at the same time it is known quite well (see for instance Bronk, 1979: 669) that the coefficient of variation for the exponential distribution is 1 (1 = c). This implies that the ratio between the coefficient of variation of the data’s values and its ranks is always 1 whenever the data are exponentially distributed. In turn this also implies that there is a general expression for that ratio, which is dependent on the value of the coefficient of variation as obtained from the data’s values, or:
From expression (7) we can see that if the coefficient of variation of the data’s values is 1 we are back to taking the ratio between expression (2) and (6), in terms of which we know that the data are exponentially distributed. On the other hand, we can see that for any distribution other than the exponential distribution, the typical value for the coefficient of variation of the data’s values is that captured by expression (6). Then by substitution of expression (6) into expression (7), followed by rearrangement, we obtain the De Vergottini inequality in terms of its limiting solution:
Outside of the typical value, it is generally known (see for instance Hürlimann, 1995: 263) that for alternative unimodal distributions—whatever they might be—the coefficient of variation of their values ranges from a minimum of zero to a maximum that is not strictly equal to 1 (0 ≤ c < 1). Therefore the corresponding one-sided non-strict representation of the De Vergottini inequality is:
For now let’s keep to expression (8). Re-entering expression (2) into expression (8), we get:
Piesch (2005: 284) refers to this numerical solution as one of the important special cases of the De Vergottini inequality. Its relevance will now become apparent in the demonstration of the Glasser inequality for the Gini index, which follows next. This latter inequality is named after the derivation method proposed by Glasser (1962b: 652–3). However the rendered mathematical treatment of this method is heavy and rather inaccessible. The intention here is to follow a simpler route, as part of keeping the algebra accessible. The starting point is the reminder by De Vergottini (1950: 453) that the Gini index is twice the area (A) of the Lorenz curve:
In expression (11), cov is the covariance between the observation values (xi) and their ranks (i), n is the number of observations and n the mean of their values. In addition, since the familiar Stuart and Pearson correlations are asymptotically equal, the covariance term in expression (11) is expressed with respect to the Stuart correlation.
Entering the standard deviation of the data’s ranks from expression (1) into expression (11), followed by manipulation and the regrouping of terms, yields the following re-expression for the Gini index:
We can see that the square root term for the number of observations in expression (12) approaches 1 as this number increases, which is approached quickly from as few as 4 observations. This reduces the Gini index to a product between the coefficient of variation of the data’s values and Stuart’s correlation with their ranks, moderated by a constant given by the inverse of the square root of 3. But we already know that the Stuart correlation has a maximum value of 1. Thus, we can replace the Stuart correlation by its maximum value, and by doing so simultaneously move to working with the ranks of the data. In effect rank transformation is being done by simply replacing the data’s values with their ranks (Conover and Iman, 1981: 124). This in turn implies rewriting expression (12) in terms of the ranks of the data:
An empirical confirmation of expression (13) is provided by Glasser (1961: 177) who shows that the approximate equality converts to a strict equality with a growing number of observations.
Piesch (2005: 264, 269) regards expression (13) as an extension of the De Vergottini inequality. This is understandable since the right hand side of the expression depicts the De Vergottini inequality by its limiting solution. Then by substitution of expression (8) into expression (13), we can rewrite the latter expression into equality between the Gini index and the coefficient of variation:
Expression (14) is a proof of the Glasser inequality for the Gini index in terms of its limiting solution. In short, as the number of observations increases, the Gini index and coefficient of variation are equal. This explains De Vergottini’s fait accompli result that as the number of observations increases the maximum value of the Gini index is one-third. We can see that the finding is not accidental. It gives the limiting value of the coefficient of variation for the observations of the data as per expression (10). Because of the equality in expression (14) the same value is then take up by the Gini index. Alternatively, it can also be obtained from the ranks of the data, by substituting expression (2) into expression (13):
We know that the square root term for the number of observations in expression (15) converges to 1 with 25 or more observations. This implies that as the number of observations increases the Gini index becomes approximately one-third. Piesch (2005: 284) refers to this last result as the other important special case of the De Vergottini inequality. While we know that the Gini index takes on values that are between zero and one, as the number of observations increases its value tends to one-third. Since the value of expression (15) is the same as that of (10), it follows that the ratio between them is 1, thereby confirming the asymptotic equality between the Gini index and the coefficient of variation. It also follows that by virtue of replicating this equality, the alternative solution of the Glasser inequality for the Gini index is also established, namely that in the case of fewer than 25 observations the coefficient of variation will systematically overstate the actual level of relative variability when compared to the Gini index. To see this note that as per expression (9), the corresponding one-sided non-strict representation of the Glasser inequality is:
As already mentioned an alternative but more demanding proof of expression (16) is done by Glasser (1962b: 652–3). This aside, the important outcome of the Glasser inequality is that it has a solution in terms of which the coefficient of variation is asymptotically equal to the Gini index. The prospects of this for extending McKay’s approximation to the Gini index are considered next.
Bronk (1979: 668–9) provides a proof for a number of well-known results concerning the coefficient of variation, namely:
When the coefficient of variation of the data is equal to zero its sampling distribution is uniform; When the coefficient of variation of the data is equal to one its sampling distribution is exponential; When the coefficient of variation of the data is anywhere between these extremes, as well as when it approaches them, its sampling distribution is indeterminate in the sense that it can take any positively skewed unimodal form.
These findings have also been reported by Hürlimann (1995: 263) and independently reproved by Hwang and Lin (2000: 135–44, 144). Consequently, the range of the coefficient of variation does not only denote abstract values. More importantly it gives signals about the shape of the data’s distribution. But the reality that the coefficient’s values fall within a range of possible values is the main reason why as summarised by Hwang and Lin (2000: 1979):
Unfortunately, the exact probability distribution of the sample coefficient of variation under most populations is still unknown.
As mentioned, when the coefficient of variation of the data is anywhere between zero and one, its sampling distribution is indeterminate in the sense that it can take any positively skewed unimodal form. This said, there is no need to despair that the exact probability distribution of the sample coefficient of variation is unknown. After all, whatever this distribution might be, its shape is positively skewed. Acting on this knowledge, McKay (1932: 697–8) proposed that the sampling distribution of the coefficient of variation can be approximated by the Chi-square distribution with n-1 degrees of freedom. Because this distribution is positively skewed it immediately lends itself as a natural contender for the sampling distribution of the coefficient of variation. On Egon Pearson’s advice (1932: 703), Fieller (1932: 699) replicated McKay’s proposed approximation, and came to the conclusion that ‘ … the approximation … is … quite adequate for any practical purpose’. Pearson (1932: 703) followed up with another independent assessment likewise reaching the same conclusion. Iglewicz and Myers (1970: 167–9) continued re-evaluating McKay’s approximation finding what McKay, Fieller and Pearson before them had already found. In their case they concluded that (Iglewicz and Myers, 1970: 169):
… the … approximation … of … McKay’s can certainly be recommended on the basis of both accuracy and simplicity.
There have been more studies that have confirmed this finding, notably those by David (1949: 388–90); Iglewicz, Myers and Howe (1968: 581); Umphrey (1983: 630–4); Reh and Scheffler (1996: 451–2); Vangel (1996: 21, 24–5); Forkman and Verrill (2008: 10–1); Forkman (2009: 234); George and Kibria (2012: 1226–34, 1239) and Gulhar et al. (2012: 48–50, 55–8, 61). The gist of these various studies is that, irrespective of the distribution of the data, McKay’s approximation is accurate with any number of observations. Sometimes this is technically described by the statement that McKay’s approximation is valid provided that the population coefficient of variation does not exceed its limiting value of one-third (Forkman, 2009: 234, 239). By recourse to the De Vergottini inequality, as per expressions (9) and (10), we can see that this condition is satisfied for any number of observations—and conclude, like Fieller, that the approximation is practically adequate for any purpose.
However the proof that the Gini index is asymptotically equal to the coefficient of variation automatically implies that McKay’s approximation extends to the Gini index too. This explains Glasser’s (1961: 177, 179–80) findings that the distributional behaviour of the Gini index coincides with that of the coefficient of variation, such that:
When its value is zero its sampling distribution is uniform; When its value is one its sampling distribution is exponential; For any values in-between, its sampling distribution is indeterminate in the sense that it too can take on any positively skewed unimodal form.
In practical situations we know that the values of the Gini index also fall within a range of possible values between its minimum and maximum limits of zero and one respectively. As a result the exact probability distribution of the sample Gini index is also unknown except for knowing that its shape is unimodal and positively-skewed (McDonald, 1981: 168–9). This again makes the Chi-square distribution a natural contender because it fulfils these shape requirements. Gerstenkorn and Gerstenkorn (2003: 470–1) report that Kamat (1953: 452; 1961: 170, 172–4) and Ramasubban (1956: 120–1; 1959: 223) have exploited this to show that, like the coefficient of variation, the sampling distribution of the Gini index is approximated by the Chi-square distribution. But neither Kamat (1953: 452; 1961: 173–4) nor Ramasubban (1956: 120–1; 1959: 223) explained why they found the Chi-square distribution to be a reliable approximation for the sampling distribution of the Gini index, except to emphasise that it is an empirical regularity. Of course from the proof of the Glasser inequality, we know that both the Gini index and the coefficient of variation are directly proportional to each other. By extension it follows that both will have the same approximate sampling distribution. Then, by default, McKay’s approximation applies to the Gini index too.
In practice, McKay’s approximation is usually carried out by McKay’s original confidence interval and/or by McKay’s modified confidence interval (George and Kibria, 2012: 1227–8; Gulhar et al., 2012: 49; Vangel, 1996: 24). Furthermore, Vangel (1996: 25) finds that with a large or growing number of observations the computed values of the original and modified McKay confidence intervals are the same up to the third digit after the decimal. The original interval is given by:
While the modified interval is given by:
By substitution of expression (14) into expression (17), the corresponding original McKay confidence interval with respect to the Gini index is:
In turn, after substituting expression (14) into expression (18), the corresponding modified McKay confidence interval with respect to the Gini index is:
In a nutshell, expressions (19) and (20) represent the extension of McKay’s approximation to the Gini index in terms of its respective confidence intervals. We can also see that there is nothing mathematically complex or computationally intensive in the intervals. They are subordinated to the well-known Chi-square distribution thought from undergraduate statistics. To remind, the lower and upper critical Chi-square (χ2) values are denoted by l and u respectively. They can be extracted from a Chi-square distribution table with n-1 degrees of freedom, where n is the number of observations, or alternatively outright computed from the familiar Wilson–Hilferty normal-based approximation (Krishnamoorthy, 2006: 161). The presence of the absolute values in the denominator of both confidence intervals prevents the possibility of cases where the limits of the intervals do not exist in their absence. A practical illustration of this possibility is provided by Wong and Wu (2002: 74, 80).
More importantly the intervals for the Gini index have two clear practical uses. First, they are a tool for determining the reliability of the estimates of the Gini index as they will indicate what their precision and bias is. We know that in a strict statistical sense accuracy is precision without bias (Grubbs, 1973: 54–6, 66). We can tell precision from the width of a confidence interval, since the width represents the largest error of estimation we are likely to make with the sample size at hand when deriving the expected or average estimate for the population value of a statistic, which in this case is that for the Gini index. We can also tell what the bias is, because bias as we know is half the width of a confidence interval.
So from a confidence interval we can comfortably establish if the expected or average Gini estimate is reliable, that is, an accurate representation for its population value. Second, the confidence intervals are a decision-making tool enabling hypothesis testing as to the magnitude of the Gini index estimates. In this way they stand to discourage decision-making with the index, which does not pay attention to the margin of error in the estimates. These practical uses are the subject of an illustration in the incoming section.
Brandolini et al. (2010: 279) reported on estimates for the Gini index of 24 countries of the European Union, as estimated in 2006. In all cases the estimates are the compiled and published official statistics by the Statistical Office of the European Union, Eurostat. They are derived from the gross earnings surveys of full-time workers as employed in 2006 in each country. Table 1 reproduces these estimates.
It is not the countries that matter in the table, but rather the striking peculiarity that the estimates for the Gini index are published without any disclosure as to their accuracy. This corroborates with the description referred to in the outset of the present article that typically the estimated Gini index is reported without its associated margin of error. However with the extended McKay confidence intervals for the Gini index, as per expressions (19) and (20), it becomes easy to revise this situation. In the present case, the computations for these expressions are done with the data in Table 1.
By way of a reminder, there are a number of conventional levels to choose from in the construction of confidence intervals, such as 90, 95 or 99 per cent confidence level. These give respectively the 90, 95 and 99 per cent confidence interval. There is nothing magical about these levels. They are just statistical conventions about the role of chance we are prepared to give in the analysis of data. As part of the illustration, the present example is done at the 95 per cent confidence level. Also by way of another reminder, given that in each case we are dealing with a large number of observations for which there are no commonly available Chi-square tables, the percentage points of the lower and upper Chi-square values of the Chi-square distribution are computed from the Wilson–Hilferty approximation. The computed estimates of the original and modified McKay confidence intervals for the Gini index are shown in Table 2.
Gini Index Estimates, Official Statistics, 2006
Gini Index Estimates, Official Statistics, 2006
As mentioned already, in practice in the case of large samples, estimation by the original or modified McKay confidence intervals does not matter much, as the computed values from both intervals are known to be the same up to the third digit after the decimal. This is also borne out in the present case, where up to the third digit the numbers from either interval are the same. This is why only one set of interval limits is given in Table 2.
Extended McKay’s Confidence Intervals for Gini Index
As an estimation devise or technique, we know that the narrower a confidence interval is, the greater its accuracy is. The results in Table 2 show that the extended McKay confidence intervals for the Gini index are in themselves reliable estimation devises, as their respective limits replicate the expected Gini estimates. Either of the intervals reveals that across the board the official Gini index estimates have high precision. Rounded off to the nearest first digit, their estimation error ranges between 1 and 2 per cent.
In turn, the Gini index estimates are accompanied by low bias. The bias in the estimation of the Gini estimates ranges between 0.5 and 1 per cent when rounded off to the nearest first digit. It is known from Chebyshev’s theorem that at least 75 per cent of observed values will lie within two standard deviations of their expected range if the values are free from bias (Richardson et al., 2004: 19). As a consequence, a statistical rule-of-thumb is to treat bias as unproblematic or trivial if it does not change an observed estimate by more than 25 per cent or 0.25 percentage points (Mooney and Duval, 1993: 33). The current bias figures are well-below this cut-off, that is, tolerance value. While this gives any public user surety in the published Gini index estimates, the point simply is that their precision and bias should be reported as part of disclosing relevant information about the published numbers. The extended McKay confidence intervals for the Gini index make this easily possible as they rely on historically well-established techniques.
An acknowledgment of the existence of bias in the estimates, also suggests that reporting the estimates with and without the impact of bias exposes their credibility in terms of the estimation processes associated with their measurement. Table 3 captures this, after applying the bias figures from Table 2 to the estimates in Table 1. As the results in Table 3 show, it is precisely because of a relatively bias-free estimation, that the published but uncorrected for bias Gini estimates can be trusted as an actual official statistic. Without this disclosure, this becomes a matter of assumption rather than fact. Because the confidence intervals are bi-directional, equally the detected bias is also bi-directional. As a result there is a lower and an upper bias-corrected estimate to consider. Fortunately because of low bias in the present case, the ratio between these estimates is 1 when rounded off to the nearest first digit, implying that the official but bias-unadjusted Gini index estimate is a reliable representation of the expected Gini index value. Of course, had this bias been higher the story would have been different. But, again if the Gini index confidence intervals are not produced this would simply be unknown.
Suppose now that after being satisfied with the credibility of the published Gini index estimates, a policy analyst is keen to know whether income inequality in the reported countries is low or high. According to a United Nations expert group (1996: Appendix 2) there is low income inequality whenever the Gini index is anywhere between 0.25 and 0.40 percentage points. The analyst wishes to know if the observed Gini index estimates adhere to this benchmark.
Gini Index Estimates with and without Bias Correction
One option to confirm the above is to use the confidence intervals for the Gini index from Table 2, and conduct a two-sided hypothesis test by setting a null-hypothesis of low income inequality at the lower bound of 0.25 percentage points for a strict test or at the upper bound of 0.40 percentage points for a lenient test. The other option, which readily flows from the confidence intervals as depicting a range of possible values, is to test for the hypothesis of low income inequality by using the confidence intervals just as they are given that their range already evaluates the possible proposition for low income inequality. In this latter case, with the presently adopted 95 per cent confidence level, the analyst simply needs to look up the range of the confidence intervals in Table 2 and compare this to the recommended benchmark range. From this the analyst will be able to conclude that the Gini index estimates support the hypothesis of low income inequality in 12 countries (Austria, Belgium, the Czech Republic, Denmark, Estonia, France, Greece, Hungary, Italy, Slovakia, Spain and the United Kingdom), that in the case of 5 of the countries (Cyprus, Lithuania, Luxembourg, Poland and Sweden) this is a borderline conclusion because the upper Gini index limit of the confidence intervals of their Gini index estimates are higher than the upper limit of the benchmark range, and that in 7 of the countries (Finland, Germany, Holland, Ireland, Latvia, Portugal and Slovenia) there is no evidence of low income inequality. There is a 95 per cent chance in the correctness of this conclusion, bar a 5 per cent possibility of being wrong. However, even then the analyst can remain confident in the reached conclusion since a similar deduction is also reachable with the bias-corrected official estimates for the Gini index from Table 3, as they too are attained at the same confidence level.
The foregoing completes the intended demonstration. Before closing it is also worthwhile to mention that the present findings have an immediate relevance to the measurement of the venerable signal-to-noise ratio in the area of quality control. Customarily this ratio is defined as the inverse of the coefficient of variation of the data’s values, and reported without confidence intervals (Noori, 1989: 323). However concerns should be harboured where such practice is observed. The issue is with the coefficient of variation itself. As Hürlimann (1998: 128) explains:
One must warn against blind application. The measure has been built starting from the variance, and there is almost general agreement that the variance is appropriate for … measurement only for normal (or approximately normal) distributions.
While Hürlimann emphasises the variance, the same can be said of the mean too, which is the other component of the coefficient of variation. As famously demonstrated by Iglewicz (1983: 404–11) the two cease to accurately estimate scale and location respectively under imperfect data conditions. This is when:
The distribution of the observations is non-symmetrical; The distribution has outliers; The number of observations is few; and lastly There is considerable variation among the observations at hand.
In the foregoing cases the choice of the measure that can accurately detect relative variability is relevant. Essentially, it is obvious from expression (16) that this is at the heart of the Glasser inequality, which highlights that if we want robustness from the signal-to-noise ratio, then we have to choose the Gini index for its estimation. This is the same as redefining this ratio as the inverse of the Gini index as per the Glasser inequality.
Acting on a seeming suggestion by Mandelbrot, this article extends McKay’s original and modified confidence intervals to the sample Gini index. Generally, the disclosure norm is to compute and report the Gini index as a point estimate only, leaving out its margin of error—and by default its expected range—due to the supposed complexity of its derivation. The introduced McKay confidence intervals for the Gini index give a viable practical alternative against such incomplete estimation and reporting. On the one hand, they make it practically possible to find the reliability of the estimates of the Gini index by computation connected with the popular Chi-square distribution. On the other hand, with recourse to the same distribution, they also act as hypothesis-testing instruments in the decisional analysis of its values. By way of an example, an illustration of these practical uses is provided to demonstrate that the extended McKay approximation for the Gini index is certainly recommendable on the basis of both accuracy and simplicity.
Footnotes
Acknowledgements
I am indebted to Professor Eon Smit for his valuable comments and suggestions. Any errors are my own. Lastly, the views expressed in this article do not necessarily represent those of Statistics South Africa.
