Abstract
Models decomposing the redistributive effect of fiscal systems into vertical and horizontal effects are extensively used by practitioners. Duclos, Jalbert, and Araar’s model, despite its advantages, has not yet been widely employed in empirical research, possibly due to a relatively challenging implementation procedure that involves the estimation of expected postfiscal incomes. To override these difficulties, the designers of the software DAD have incorporated a module for implementation of the model. However, the application of this module on Croatian tax-benefit system data revealed certain inaccuracies in the results. Carefully unfolding the calculation and estimation procedures needed for implementation of the model, this article instructs practitioners on how to correctly apply the model and helps DAD designers improve their module.
Duclos, Jalbert, and Araar (2003; henceforth DJA) have designed a comprehensive model to decompose the redistributive effect (RE) of a fiscal system into vertical, classical horizontal inequity (henceforth CHI), and reranking effects. The model is built into the framework of the Atkinson-Gini social welfare function (henceforth AGF), which first converts incomes into utilities employing the Atkinson’s (1970) utility function, and then aggregates utilities using rank-dependent weights, which underlie the S-Gini and S-concentration coefficients proposed by Donaldson and Weymark (1980) and Yitzhaki (1983). 1
The DJA model has certain advantages over its competitors, the widely acknowledged Kakwani’s (1984; henceforth K84) and the Aronson, Johnson, and Lambert’s (1994; henceforth AJL) decompositions of RE. To measure the CHI effect, the researcher must determine the set of counterfactual CHI-free or expected postfiscal incomes (EPIs). While the AJL model relies on the formation of arbitrary groups of close equals in this task, the DJA model employs purposefully designed statistical procedures. Consequently, the implementation of the DJA model requires a certain expertise related to data smoothing and curve-fitting methods. To facilitate the application of the DJA model in empirical analysis, a module for calculation of the DJA indices from the sample data is incorporated into the software DAD (Duclos, Araar, and Fortin 2010; henceforth DAD-DJA).
The use of DAD-DJA in research on the Croatian tax-benefit system revealed certain inaccuracies in the results. Specifically, when the ethical parameter of AGF is set to zero, the CHI effect in the DJA model should be equal to zero by construction. However, the estimated value of the CHI effect was significantly different from zero. Analysis has shown that DAD-DJA produces upward biased estimates of EPIs in the low prefiscal income region. Furthermore, it was revealed that the fitting procedure in DAD-DJA contains a “bug,” producing unreasonably high estimates of EPIs for the top prefiscal income units in the sample.
In an attempt to obtain fully accurate estimates of DJA indices, independent procedures have been developed. They are thoroughly explained in this article to assist practitioners in implementing the DJA model and to help DAD designers improve the working of DAD-DJA. A brief overview of data smoothing methods is provided, accompanied by advice on how to accurately obtain EPIs estimates. Relationships with other measurement models are explained.
The rest of the article is organized as follows: the DJA Model section briefly exposes the elements of the DJA model and its connections with other decompositions. The Calculation of Indices section extensively describes the procedures of data preparation, estimation, and calculation of various elements of the DJA model, and employs them on a simple hypothetical population of four income units. In Application: Croatian Tax-Benefit System section, the procedures are applied to data on the Croatian tax-benefit system, and the results are compared with those obtained by DAD-DJA. The conclusions are presented in the final section.
The DJA Model
Postfiscal income is equal to prefiscal income minus taxes plus benefits. RE is the change of income inequality induced by a fiscal system consisting of taxes and benefits. In measurement terms, we have that
In the DJA model, inequality indices
The DJA model decomposes RE as follows:
The vertical effect,
In equation (3),
In the special case where
It can be shown that
Consequently,
In another special case, where
Calculation of Indices
Data Preparation
A typical research uses the following data for a household or family i: (a) unequivalized pre- and postfiscal incomes,
We form the
To obtain the matrix
From
The sample estimates of quantiles p and the weights
where
When a large group of prefiscal exact equals exists in the sample, one of the inequality indices would be biased if based on the weights
Thus, the original weights
Alternatively, we could use the original weights and randomize the order of income units within each group of exact prefiscal equals. This procedure would reduce the bias to an insignificant level, but each possible ordering of income units would still result in different values of
Finally, analogous to the preceding procedures, the estimates
Indices of Inequality
The following equations show how to obtain utilities, the Gini-Atkinson welfare index, and the inequality index for prefiscal incomes
where
To obtain the sample estimates of EPUs,
However, the following identity says that the whole procedure of estimating EPUs can be circumvented, saving the practitioner’s time and energy in sensitivity analysis using multiple scenarios for
To understand why (11) holds, recall that
Observe that according to (A3) we would obtain the identical result for
Estimation of EPIs and Utilities
Unlike the estimation of EPUs, the evaluation of EPIs cannot be avoided. To obtain the sample estimates of
The estimation of EPIs represents the greatest challenge in the implementation of the DJA model. Although parametric models (such as polynomial regression) can be appropriate for some data sets, it is better to rely on nonparametric approaches, assuming no a priori functional relationship between post- and prefiscal incomes. One such approach is the “kernel-weighted local polynomial regression” (KWLPR). A description of the method can be found in Fan and Gijbels (1996), Wand and Jones (1995), Keele (2008), and Härdle (1990), while the software applications include Stata 12 (function lcpoly), R (function loess, package lokern, etc.), and XploRe (function lpregxest).
The choice of the degree of polynomial (p), the type of the kernel function, and the size of the kernel half-bandwidth rests on the analyst. For
Another interesting smoothing technique came to light during the research: the “Fourier series in trigonometric form” (FSTF), which is a sum of the sine and cosine functions describing a periodic signal (Faunt and Johnson 1992). The estimation procedure is programmed in Matlab R2011b’s Curve Fitting Toolbox 3.2, which contains several other fitting methods, such as smoothing splines.
DAD-DJA and supporting documentation 5 do not inform us which fitting method is used to obtain EPIs for estimation of DJA indices. However, DAD incorporates a separate module, “Non Parametric Regression,” enabling us to estimate EPIs independently of DAD-DJA. Two basic methods are offered: NWE and LLE (henceforth, DAD-NWE and DAD-LLE). Experimentation with different options and choices offered by the module revealed that in estimating EPIs DAD-DJA in fact employs DAD-LLE, using the default set of parameters.
Before moving further, we offer the following advice to help judge whether the estimates
Although the fitting methods and their software implementations ensure optimality in the statistical sense, the analyst still has the freedom and the responsibility to change some of the parameters or the whole estimation method if the results contradict her or his knowledge of the appropriate shape of the EPIs curve. An example is a too “wiggly” curve, in which case we have to “stretch” it, perhaps by raising the kernel half-bandwidth. Another example may be the existence of certain kinks or local minimums (maximums) we are aware of, which are not reflected by the EPIs estimate.
For certain data points, the programmed fitting procedures may produce irregular results. Some software tools are “smart” in such cases, leaving a blank space instead of the estimate, while others are not. Anyway, if this happens, we should fill in the corresponding estimate manually, using the best-guess approach.
A simple preliminary test of the correctness of the approximation
The discussion in The DJA Model section indicated that when
From equation (13) follows another test: the inequality indices
Decompositions
Having defined all the indices needed, we can present RE and its decompositions in terms of sample estimate formulas. RE is obtained as
where the last row in equation (14) arrives from the property (11), by which
Setting
Simple Hypothetical Example
We return to the example of four hypothetical households from The DJA Model section to illustrate how the DJA model implementation procedures work. There are two groups of prefiscal equals in the sample: A and B with prefiscal income of 10 each belong to the lower quantile, whereas C and D with prefiscal income of 20 each belong to the upper quantile of prefiscal income distribution.
The first column in table 1 shows the “original” weights
Hypothetical Population: Weight, Incomes, and Utilities
Note. Weights are obtained for
As equation (11) explains, the estimate of
According to (A5), for
and
which is identical to the result that would be obtained by (A3):
Thus, our hypothetical example confirms the identity (11). On the other hand, indices based on the ‘wrong’ weights,
and
The estimate of
and
All inequality indices for
Indices Obtained for Hypothetical Population
Another set of CHI and reranking effects is derived using
Recall that equation (13) says that the inequality indices based on
Finally, we look at how DAD-DJA deals with this small hypothetical case. The DAD supporting documentation tells us that the estimate of
Application: Croatian Tax-Benefit System
Data
We analyze the fiscal system consisting of social security contributions (SSC) for the pension, health, and unemployment insurance funds, personal income tax and surtax (PITS), public pensions, and cash social benefits. 6 The data on incomes come from the Croatian household budget survey (Anketa o potrošnji kućanstava; APK) for 2008, whose sample contains 3,108 households. Since APK registers only net incomes of household members, the amounts of prefiscal income, PITS, and SSC are obtained by a microsimulation model.
Postfiscal income of household i is obtained as
Before analyzing the results of the DJA decomposition, we observe the features of the data set. The dots in the scattergram (figure 1) are the post-fiscal and prefiscal incomes of sample income units, expressed in terms of the mean prefiscal income (mpfi). The full line shows EPIs obtained by KWLPR (see the next section for details on estimation). The dotted line represents the cumulative density, which tells us, for each prefiscal income X, the proportion of all income units having prefiscal income below X (on the right axis). We can observe that quite a large proportion of units, about 7 percent, have zero prefiscal income (group A), while the next 13 percent of units have prefiscal income below 10 percent of mpfi (group B).

Scattergram of pre- and postfiscal incomes
The mean postfiscal incomes of groups A and B are 64 percent and 54 percent of mpfi, respectively. Observe that the EPIs curve is decreasing on the interval [0, 0.1]. The following three facts taken together can explain the curious feature that the mean postfiscal income is decreasing. First, for the majority of pensioners’ households, a public pension is the only source of income; since public pensions are benefits in the current scenario, the pre-fiscal income of most pensioners’ households is zero. Second, majority of households with zero prefiscal income (group A) are pensioners’ households. Third, pensions are on average higher than other social benefits.
Estimation of EPIs and the Decomposition
The indices of the DJA decomposition are estimated by three models, using three different fitting methods described in Estimation of EPIs and Utilities subsection.
In model A, EPIs are estimated by KWLPR programmed in Stata 12. Following Bilger (2008), we use the third-degree local polynomials, employing the Epanechnikov kernel. The optimal half-bandwidth of the kernel obtained by the program was equal to 6.7 percent of mpfi, and it was increased by one half. In model B, EPIs are obtained using FSTF programmed in Matlab R2011b’s Curve Fitting Toolbox 3.2. The number of harmonics is set to 7; the “Trust-Region” algorithm is employed with the robust fitting option turned off. In both models, the top five prefiscal income units are excluded from the fitting process, and their values of
The aim of model C is to replicate the results obtained by DAD-DJA. We employ DAD-LLE to estimate EPIs, with all observations included in the fitting process. To estimate
Before moving on to the results, let us look at the shapes of the different EPIs curves, shown in figure 2, concentrating first on the bottom part of the income distribution. While both A and B reflect the initial fall in EPI, discussed earlier, C does not, that is, its EPIs curve is rather flat on the whole interval. For prefiscal income of zero all estimates are roughly the same, but in the prefiscal income interval [0.025, 0.42] of mpfi, C’s EPIs lie above those estimated by A and B. On the prefiscal income interval [0, 0.5] of mpfi, the mean of EPIs obtained by A (B) is 0.5764 (0.5775) of mpfi, which is very close to the mean postfiscal income for actual values, equal to 0.5762. On the other hand, the mean of EPIs obtained by C is 0.5881, or 2 percent above the actual mean. This suggests that C overestimates EPIs for the lowest incomes. For prefiscal incomes above 0.5 of mpfi, the EPIs of B and C are almost identical, while the EPIs curve of A is “more flexible” and intertwining the other two curves.

Scattergram of pre- and postfiscal incomes
Models A and B convincingly pass the test from equation (12), as the ratios
Decomposition of Redistributive Effect for
The estimates
Model C thus underestimates the vertical effect by about 3 percent of RE for
Decomposition of Redistributive Effect for Different Combinations of
Conclusion
Models decomposing the RE of fiscal systems into vertical and horizontal effects are extensively used by practitioners. The DJA (2003) model, despite its advantages over some other models, such as the Kakwani’s (1984) and the Aronson, Johnson and Lambert’s (1994) decompositions of RE, has not yet been broadly employed in empirical research. The reason may be the relatively complex implementation procedure, which involves nonparametric methods in estimation of EPIs.
To override these estimation and calculation difficulties, the designers of the software DAD have incorporated a module for estimation of the DJA model indices, here referred to as DAD-DJA. However, as the application data on the Croatian tax-benefit system indicates, DAD-DJA produces somewhat inaccurate estimates of EPIs, resulting in biased values of DJA model indices. This article carefully explains the estimation procedures needed to obtain the indices of the DJA model, and the problems occurring in DAD-DJA implementation.
The estimates of EPIs are obtained by two fitting methods: kernel-weighted local polynomial regression and Fourier series in trigonometric form. Both achieve reasonable fit of the data at stake, unlike the method built into DAD-DJA, which seems to overestimate EPIs at the bottom region of prefiscal income distribution. Furthermore, we have realized that the fitting procedure in DAD-DJA contains a “bug,” producing unreasonably high estimates of EPIs for the top prefiscal income units in the sample.
We have shown how the estimation of EPUs can be circumvented, saving a practitioner time when doing multiple-scenario analysis. Instead of estimating EPUs for each different value of parameter
Footnotes
Appendix
Acknowledgment
The author would like to thank two anonymous referees, Jean-Yves Duclos and Slavko Bezeredi, for their very useful comments and suggestions.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
