Abstract
Expert surveys provide a standardized way to access and synthesize specialized knowledge, thereby, enabling the analysis of a diverse range of concepts and contexts that might otherwise be difficult to approach systematically. However, while studies of public opinion have long argued that cognitive biases represent potential problems when it comes to the general population, less attention has been paid to similar issues among expert respondents. This study examines one form of cognitive bias, hindsight bias. Hindsight bias refers to the tendency to retrospectively exaggerate one’s foresight of a particular event. We argue that hindsight bias is a potential problem when it comes to retrospective evaluation due to the difficulty involved in separating our assessments of the pre-crisis period from the knowledge that a crisis occurred. Using disaggregated data from the Varieties of Democracy Project, we look for evidence of hindsight bias in coders’ evaluations of the periods that preceded major crises of democracy. We find that coder disagreement is significantly higher in pre-crisis scenarios than in our control group. Concerningly, despite this disagreement, coders remain similarly confident in their assessments. This represents a potential problem for those who seek to use these data to study democratic breakdowns and transitions.
Introduction
Expert surveys are an increasingly popular method of studying a range of political phenomena. Yet, despite an appearance of authority, expert surveys remain vulnerable to many problems associated with lay questionnaires. While studies of public opinion have long explored cognitive biases among the general population (see Kuklinski and Quirk, 2000), less attention has been paid to bias among experts. Presumably, this is because experts are better positioned to offer more objective (or, at least, less overtly biased) judgements. The question remains: To what degree does expertise shield respondents against errors in reasoning?
This study examines one form of cognitive bias: hindsight bias. Hindsight bias refers to the tendency to retrospectively exaggerate one’s foresight of an event – a problem that is well-documented in other areas of political science (see Lebow, 2009). We argue that hindsight bias is particularly insidious when it comes to retrospective evaluations of crises due to the difficulty involved in separating our evaluations of the pre-crisis period from the knowledge that a crisis occurred. Complicating matters further, political memory is often ideologically polarized, as interpretations of past events may be embedded in contemporary debates that (re)frame past conflicts.
Until recently, identifying hindsight bias in expert surveys has been difficult, as many indices that incorporate expert judgements only provide aggregate results. While there are valid reasons for withholding disaggregated responses, without these data, we cannot fully discern the extent to which experts (dis)agree, especially in the absence of a corresponding error term.
We look for evidence of hindsight bias in data from the Varieties of Democracy (V-Dem) Project, which relies on over 3000 anonymous experts (Coppedge et al., 2018). We examine coders’ evaluations of two groups of democracies: one in which a crisis occurred and one where no evidence of a crisis was observed.
Our findings suggest two areas of concern. First, we find that disagreement is significantly higher in the pre-crisis cases than in the corresponding control group for almost every indicator. Second, despite this disagreement, coders remain confident in their assessments. This poses a problem for those studying democratic breakdowns, as our findings not only suggest that coders’ knowledge of a crisis colours their evaluations, but also that coders may not be fully aware of this.
This article proceeds as follows. The next section introduces the cognitive phenomenon of hindsight bias and the expands on the relationship between memory and the measurement of democracy. This is followed by a discussion of the hypothesis, research design, and data. We then present our results and offer some interpretive discussion and recommendations, followed by a conclusion that considers the applicability of our findings to a wider research agenda interested in questions of democratization and democratic decline.
Democracy, measurement and memory
Democracy is a contested concept. There is no consensus among those who study democratic attainment as to which factors should be included. This fact is implicitly acknowledged in the V-Dem data, which permit researchers to choose between five democratic indices or construct their own measures from a diverse menu of indicators. While this study does not endorse a particular definition of democracy, it is necessary to acknowledge this disagreement in order to understand the different ways in which expert surveys have operationalized the concept of democratic attainment, as well as the different ways in which coders are likely to understand this concept.
Expert surveys
While we might assume that experts are more likely to produce reliable judgements than lay people, ‘expert’ respondents might have imperfect information; they might understand concepts and metrics differently; and their judgements may be affected by ideological convictions (Maestas et al., 2014; Martinez i Coma and Van Ham, 2015). Expert surveys are therefore not immune to fundamental problems of measurement error. In particular, the literature emphasizes two areas of concern: inter-coder reliability and data aggregation.
First, because experts may have different understandings of relevant concepts, they may assign different values to the same case. This problem, known as differential item functioning (DIF), refers to variation in how experts apply conceptual tools such as ordinal scales to evaluate cases. While some discrepancy is anticipated, respondents’ judgements should be similarly calibrated. Numerous measures have been developed to assess this (see Hayes and Krippendorff, 2007; Steenbergen, 2000; Steenbergen and Marks, 2007). Yet, even though reporting at least one measure of inter-coder reliability is now considered ‘best practice’ (see Maestas, 2016), many indices do not report this or other measures of uncertainty (see Coppedge et al., 2011: 251). This remains true of widely used tools based on expert questionnaires, such as the Freedom House (2018) global rankings of freedom and democracy.
Second, the process of aggregating expert assessments adds another layer of complexity. While this issue has been discussed extensively in the literature (see Jones and Norrander, 1996; O’Brien, 1990), of particular relevance to the proceeding discussion is the way in which indices incorporate experts’ confidence in their judgements into the scoring system.
V-Dem: Measurement and aggregation
There are several ways in which V-Dem represents a major advancement in overcoming problems of subjectivity and inter-coder reliability.
First, its organizers have paid considerable attention to conceptual clarity and standardization. Consider, for example, the coder instructions for the indicator measuring freedom of discussion among men, reproduced in Supplemental Appendix A. The detailed question, two-paragraph clarification, and specific responses all address the DIF problem (i.e. that experts’ idiosyncratic perceptions affect their application of the scale described in the V-Dem codebook). While some subjectivity remains, the use of detailed, question-specific responses – as opposed to a Likert-type scale with vague categories ranging from ‘strongly disagree’ to ‘strongly agree’ – greatly reduces this subjectivity.
Second, the use of ‘bridge’ coders enhances cross-country and intertemporal comparability. About 15% of V-Dem coders are bridge coders, who code at least two different countries according to the same criterion for the same period. This allows V-Dem’s measurement model to incorporate information about coders’ assessments of countries with different historical trajectories. 1
Third, when it comes to data aggregation, V-Dem’s Bayesian item response model is sensitive to the reality that coders will have different understandings of terms like ‘somewhat’ and ‘mostly’ when inputting responses. Importantly, the aggregate data (and error terms) incorporate information about coders’ confidence in their assessments, as well as patterns of cross-coder (dis)agreement to aid in estimating variations in reliability and possible bias (Coppedge et al., 2018: 24–29).
These advancements notwithstanding, however, the remainder of this article argues that V-Dem’s experts are not entirely exempt from the possible effects of hindsight bias.
Memory and retrospective evaluation
Experts in the present may have reason to be optimistic (or pessimistic) about the future, but they cannot know for certain what will happen. Conversely, coders evaluating the past cannot separate their assessments of a regime from the knowledge that it ultimately flourished or failed. Thus, although retrospective assessment is critical to datasets such as V-Dem, it also introduces potential problems.
Hindsight bias refers to the tendency to retrospectively exaggerate one’s foresight of an event – in this case, a crisis of democracy. For Fischoff (1975), hindsight bias is motivated by a predisposition to perceive an innate necessity in the order of events that led to some significant outcome. This, in turn, is likely to affect the way individuals recall the circumstances preceding the event. In general, hindsight bias leads researchers to come to regard past events as over-determined, while the future remains highly contingent (Hawkins and Hastie, 1990).
Although the potential perils of hindsight bias have been well-documented (see Tetlock, 2005), our training as scholars leaves us particularly vulnerable to this way of seeing the world. As Lebow (2009: 59) observes, ‘social scientists make reputations for themselves by proposing new explanations or theories to account for major events’ such as democratic crises. We are, therefore, professionally predisposed to explain events as complex causal chains rather than the result of chance. This is likely to lead at least some experts to view pre-crisis contexts in light of the crisis that followed, making their assessments potentially less reliable – even if coders themselves are not fully conscious of this.
Consider, for example, the case of the Weimar Republic. No expert could possibly be unaware of this democratic breakdown or the horrors that followed. The question, then, is to what degree present-day evaluations may be tainted by this knowledge? Can we assess the quality of the Weimar Republic independent from the knowledge that it broke down? 2
Related questions arise in relation to other historical crises. In Chile, for example, the 1973 coup d’état in which President Salvador Allende was overthrown by the armed forces remains a watershed moment in the country’s history; as a result, competing representations of the pre-coup period are deeply embedded in contemporary political debates. Right-wing parties, in particular the Independent Democratic Union (UDI), which supported the continuation of authoritarian rule during the 1988 plebiscite, portray Allende’s left-leaning government as having been on the verge of plunging Chile into a communist dictatorship. This interpretation effectively recasts the dictatorship of General Augusto Pinochet as an authoritarian caretaker regime that sought to prevent a communist takeover until Chile was ready to transition (back) to democracy. By contrast, this account is utterly rejected by left-leaning parties, who depict Pinochet’s repressive regime as having illegally usurped power from a democratically elected leader. In this way, memory of the pre-coup period continues to be extremely relevant to contemporary political debates, in which the role of the armed forces and the nature of the constitution (implemented under Pinochet) remain contested.
As the Chilean case illustrates, although ideological bias differs from hindsight bias, the two problems are not entirely independent. While we cannot alter past events, the way in which these events are (re-)interpreted in the present, and their significance in relation to events that are currently unfolding, is constantly evolving. We argue that crises of democracy, by their very nature, are particularly vulnerable to shifting/competing interpretations in this way. Thus, to the extent that competing interpretations are influenced by ideological claims relevant to the present, it seems probable that hindsight bias may lead partisan coders to view the pre-crisis period not only as over-determined, but also through an ideological lens.
To be clear, we do not suggest that all expert evaluations suffer from hindsight bias, only that, as academics, we remain potentially vulnerable to its effects. It seems probable that hindsight bias may not be experienced equally by all coders; for instance, it may be that those with direct personal experience of a particular time are better positioned to assess it more ‘objectively’ than those whose knowledge is based on secondhand accounts (which may themselves be vulnerable to hindsight bias, thereby magnifying the problem). Hence, the effects of hindsight bias may also appear magnified over time, as individuals’ memories of the pre-crisis period are gradually overshadowed by knowledge of the outcome (Blank et al., 2003) and increasingly informed by secondary accounts – accounts which may not be politically ‘neutral’, as the Chilean case shows. 3
Importantly, memory experiments show that hindsight bias also has a secondary, more insidious effect, in that it affects not only participants’ actual evaluations, but also increases experts’ confidence in their retrospective predictions (Synodinos, 1986). Hence, despite their expertise (or perhaps because of it), expert coders are also likely to be unreliable judges of the quality of their assessments.
Research design
Because the V-Dem data include very little information about coders themselves – including partisan affiliations – it is not possible to test for the effects factors like ideology or familiarity with the case. 4 Thus, instead of an expert-focused analysis of variance (see Frank and Martinez i Coma, 2017; Martinez i Coma and Van Ham, 2015; Silva and Littvay, 2019), we consider context-related factors.
Hypotheses
To look for evidence of hindsight bias, we examine the level of (dis)agreement among V-Dem’s expert coders. Although one might assume that hindsight bias would lead experts to evaluate a crisis in the same (over-determined) way, thereby minimizing disagreement, we anticipate the opposite effect. This is because we do not necessarily expect all coders to share the same view of the underlying cause of the crisis – nor even to agree on whether a crisis occurred at all. (This latter point is taken up below.) Our understanding of hindsight bias – as a general tendency to retrospectively consider the occurrence of an event as over-determined – therefore emphasizes the fact that evaluators’ knowledge of an event may influence their evaluation of the period that preceded it, but we do not assume that all coders view the event as over-determined for the same reasons. Because of the inherently political nature of democratic crises, instead of a convergence of expert views, we expect hindsight bias to have a potentially polarizing effect where crises are concerned, leading experts to disagree more than they would when coding non-crisis cases.
Although the data do not permit us to test for the effects of ideological bias directly, we can measure the degree of variance among expert opinions. To test this, we construct two groups of democracies for comparison: one in which a crisis occurred, and one where no evidence of a crisis was observed. In general, we expect the level of disagreement to be low but non-zero, since experts will have slightly different understandings of questions and measurement tools (the DIF problem). More specifically, we expect disagreement to be significantly greater in ‘pre-crisis’ democracies, independent of any ‘background noise’ caused by different interpretations of the response categories.
This is in contrast to the null hypothesis, which suggests that coders’ evaluations of pre-crisis and non-crisis democracies should not differ – that is, both contexts should be similarly vulnerable to DIF. In that sense, the null hypothesis is not that coders should never disagree, but that coder disagreement should be comparable in both pre-crisis and non-crisis democracies, since the measurement tools remain the same.
Data
Our data come from the disaggregated assessments of the Varieties of Democracy Project (Coppedge et al., 2017), which covers 177 countries between 1900 (or independence) and 2016. Our focus is the electoral democracy index, which ranges from 0 to 1 and mirrors Dahl’s (1971) concept of polyarchy. It comprised sub-indices measuring the degree of freedom of association, expression, suffrage, and executive elections. These sub-indices, in turn, comprised multiple lower level indicators (see Supplemental Appendix D).
Some indicators are based on factual information gleaned from official documents. There is no variance to report in these instances. The remainder – which forms the basis of our analysis – consists of ‘subjective assessments on topics like political practices and compliance with de jure rules’ (V-Dem Institute, 2018). Approximately five coders assess each indicator, although that number varies by country, year, and indicator. The relative weight of coders’ evaluations is determined by the Bayesian item response theory (IRT) measurement model, 5 which incorporates data about coders’ confidence in their evaluations and the degree to which coders’ assessments are consistent with those of their peers.
Although V-Dem is less susceptible to many criticisms levelled against other expert surveys, the proceeding analysis suggests its coders remain vulnerable to cognitive bias that arises from retrospective evaluation.
(Pre-)crisis groups
The fact that different democratic indices reflect different understandings of democracy is particularly relevant when examining crises; what constitutes a crisis depends on how democratic attainment is operationalized. Hence, in defining a ‘democratic crisis’, our intent is to identify periods of pronounced decline in previously democratic polities. We leave the term ‘crisis’ deliberately broad in order to capture a range of scenarios. Thus, our ‘pre-crisis group’ is, in fact, five distinct groups, each based on different conceptions of ‘democracy’ and ‘crisis’. The unifying theme is that, in each case, polities experienced a steep decline in democratic attainment from 1 year to the next. 6
Table 1 provides an overview of five conceptualizations used in the proceeding analysis. The first group is characterized by a decline on V-Dem’s electoral democracy index. However, to avoid problems of endogeneity, the remaining groups are determined using different indices.
Crisis and control group definitions.
The first group is defined by a drop in a country’s V-Dem electoral democracy index of 0.25 or more from 1 year to the next. Because V-Dem does not assign category labels, there are no cut points that distinguish between democracy and authoritarianism – polities exist on a spectrum ranging between 0 and 1. Hence, while a decline of 0.25 is somewhat arbitrary, it is meant to capture both the magnitude of the crisis and the fact that such a decline is only possible where a country’s score was already sufficiently strong as to permit it. 7
The second group is defined by a drop in a country’s Polity IV combined score from +6 or higher to −6 or lower from 1 year to the next. These values correspond with Polity IV’s data manual, which distinguishes between autocracies (−10 to −6), anocracies (–5 to +5), and democracies (+6 to +10) (Marshall et al., 2002). This definition captures a crisis that would see a democratic polity collapse into an autocracy within the span of a single year. 8
Three further groups are also included for comparison. Group 3 is based on the ‘democratic breakdown’ variable in Boix et al. (2013) which identifies a transition from democracy to non-democracy. Group 4 includes polities that were coded by Cheibub et al. (2010) as democratic in 1 year but not in the next. The final group includes cases that were previously coded as democratic by either Polity IV, Boix et al. (2013), or Cheibub et al. (2010), and where at least one coup d’état (as coded by Przeworski et al., 2013) occurred in the specified year.
Importantly, the selection criteria emphasize the sudden nature of the crisis, as opposed to gradual decay. While this excludes cases in which democracy eroded slowly, the narrower focus is meant to ensure that (a) a ‘crisis’ occurred and (b) the polity in question was still ‘democratic’ in the year before the crisis.
Supplemental Appendix B illustrates which cases (countries and years) are identified by each of the aforementioned definitions. Although these indicators have been widely used in studies of regime transitions and democratic attainment, they disagree as to what constitutes a crisis of democracy. This is evident in Table 2, which reports the proportion of cases that each group has in common, as well as the number of crises this represents (in parentheses). For example, despite the fact that the V-Dem and Polity IV data are, in general, strongly positively correlated, and despite identifying a similar number of crises (17 and 20, respectively), these groups have just three crises in common. In fact, all five groups agree on just two cases: Argentina (1976) and Chile (1973). 9 Out of the 105 crises identified by the different measures, 36 are captured by only one definition. 10
Common crises.
This finding is surprising given that these definitions are intended to rule out all but the most unambiguous cases. This underscores the necessity of considering different conceptualizations of ‘democratic crisis’ in the analysis that follows, but it also highlights the degree to which experts disagree when it comes to (pre-)crisis democracies.
This disagreement has specific implications for hindsight bias: whether or not an expert believes that a crisis occurred is likely to affect their judgement regarding the pre-crisis period; since experts disagree as to what constitutes a crisis, it follows that their evaluations of the period that preceded that (non-)crisis may also be differently affected by possible hindsight bias.
In what follows, we are interested in coder disagreement in the pre-crisis periods – that is, the year that preceded the crises outlined in Table 1. While this period is somewhat arbitrary in the sense that democratic crises may have roots that stretch back many years, our emphasis on the sudden nature of the crisis allows for standardization and ensures that the quality of democracy (pre-crisis) was reasonably robust. The pre-crisis period is also critical from the point of view of those interested in identifying potential signs of trouble in the lead-up to a crisis.
Control group
The use of a control group provides a sense of what level of disagreement is typical among V-Dem experts’ assessments of different phenomena. This is important because the null hypothesis suggests that disagreement across comparable democratic cases should be similar but non-zero.
The control group is constructed using inverted interpretations of the crisis definitions above. To be included, a case must (a) have been coded as democratic by all the definitions above (except V-Dem) for the past 5 years 11 and (b) also not be identified as a pre-crisis year according to any of the definitions above (including V-Dem). This results in a pool of 1468 country-years with electoral democracy index scores ranging from 0.22 (El Salvador, 1989) to 0.91 (Sweden, 1990).
Because this number of observations greatly exceeds that of the pre-crisis groups, we use a random sub-sample for comparison. 12 This prevents artificial deflation of the standard deviation among coders’ assessments in the control group. The number of randomly selected country-years roughly corresponds with the number of pre-crisis years, however, the actual number of coder assessments varies based on the number of coders assigned to evaluate each country/year/indicator. Supplemental Appendix C lists the cases (country and year) in the control sub-sample.
Method of comparison
Examining coders’ raw assessments, we conduct two complementary analyses. First, following Silva and Littvay (2019), we attempt a multigroup confirmatory factor analysis to test for measurement invariance (MI) between the pre-crisis and control groups. Conceptually, MI is similar to DIF, in that looks at differing response patterns from coders (see Kline, 2016: 398). In the case of the V-Dem response tool, MI would occur if two expert respondents who share the same level of a latent construct both provide the same response to a question measuring it, irrespective of whether they are coding a pre-crisis or non-crisis democracy.
Second, we compare the level of coder disagreement in the pre-crisis group versus the control group using Levene’s test. 13 To be clear, we do not compare actual levels of democratic attainment; rather, we are looking at differences in the level of variation among coders’ raw assessments. We repeat this for all components of V-Dem’s electoral democracy index that are assessed by experts – 23 variables in all.
Findings
Measurement invariance
Multigroup confirmatory factor analysis allows us to test whether some model parameters differ between the pre-crisis and control groups described earlier. If the V-Dem response tool is invariant across groups, freely estimated models should be similar for both the pre-crisis and control groups. We test for invariance separately in each of the sub-indices used to construct the electoral democracy index that rely on expert evaluations: the freedom of expression index (9 indicators); the freedom of association index (6 indicators); and the clean elections index (8 indicators). 14 Modelling each lower level index as a latent variable, we run the analyses on the coders’ original responses using a simple ordinary least squares (OLS) model. We then use a likelihood-ratio test to compare the model with all parameters constrained against the model with all parameters estimated distinctly for the pre-crisis and control groups. 15 We repeat the test for each of the five distinct pre-crisis groups described earlier.
The results, reported in Table 3, indicate that for all sub-indices and pre-crisis groups, the model with distinct parameters fits better than the model with all parameters constrained. In other words, measurement of the latent concept (freedom of expression, association, or clean elections) differs between the pre-crisis and control groups, meaning that experts coding pre-crisis democracies have a fundamentally different understanding of what it is that they are measuring as opposed to coders of non-crisis cases.
Test for group invariance (comparison of pre-crisis with sample group).
Measuring variance
Table 4 reports the difference between the standard deviations of the control groups versus the pre-crisis groups. A value of 1.0, for example, indicates that there is no substantive difference in standard deviation, while a value of 1.5 indicates that standard deviation is 50% higher in the pre-crisis group than in the control group.
Percentage difference in variance.
Results are reported for the mean (W0); Supplemental Appendix E contains results for the median (W50), and the 10% trimmed mean (W10).
Values represent the difference in variance (%) between the pre-crisis and control groups; Missing values (.) indicate that there was no disagreement whatsoever among coders in the control group, resulting in a standard deviation of 0 for that group.
p < 0.05; **p < 0.01; ***p < 0.001.
CSO: Civil society organizations; EMB: Electoral management bodies
Table 4 reveals that for almost every indicator, variance is higher in the pre-crisis period than in the control group. On average, variance is between 43% and 76% higher in the pre-crisis groups when compared to the control groups, with this difference being highest in the V-Dem group.
There are several exceptions to this trend. There is less coder disagreement when it comes to the harassment of journalists and media self-censorship. This may be because these are also some of the indicators for which standard deviation is typically highest, even among the control groups. There also appears to be comparatively less disagreement in relation to some of the components of the clean elections index, especially in the Boix et al. (2013), Cheibub et al. (2010), and Przeworski et al. (2013) groups.
The starkest disagreement occurs in coders’ evaluations of the components of the freedom of association index (party ban, barriers to parties, opposition party autonomy, and the relative freedom of civil society organizations). On some of these indicators, standard deviation in the Polity IV and V-Dem groups is two to three times higher than in the control group. Taken together with the results of the multigroup confirmatory factor analysis, this underscores the finding that experts coding pre-crisis democracies have a fundamentally different understanding of what they are measuring than coders of non-crisis democracies.
Discussion
Coder confidence
Our findings raise an important question: is the higher level of disagreement in the pre-crisis contexts a result of expert coders being unsure of their evaluations? Given the difficulty of retrospective coding discussed earlier, this is a possibility. Because V-Dem asks coders to provide a corresponding measure of confidence in their assessments, this explanation is also relatively easy to consider.
The difference in coder confidence between the pre-crisis and control groups is clearest in relation to some of the indicators that have the highest levels of variance, such as the existence of barriers to political parties. However, this difference is typically modest (3% to 10%, on average), and the difference between the pre-crisis and control groups in relation to a single indicator remains lower than the average variation in coder confidence across indicators but within the same group (15% to 22%). Overall, coders in all groups were similarly confident, even though their scores reveal far greater disagreement among pre-crisis coders than among coders of the control groups. 16 If coders in the pre-crisis group were more unsure about their evaluations or struggled more with the application of the response tool, this is not reflected in their self-reported confidence.
Because familiarity with the case in question is not measured for most V-Dem coders, it is not clear how familiarity relates to confidence. Nevertheless, we can examine the confidence of bridge coders, who have been asked to code cases outside their immediate expertise. Among bridge coders of pre-crisis contexts, 17 there is virtually no difference in relation to other coders; if anything, bridge coders were, on average, about 2% more confident in their responses (78% vs 80%). This suggests that coder confidence is not necessarily a good indication of familiarity.
Although pre-crisis coders were typically quite confident, they disagreed in their actual evaluations, as shown in Table 4. Taken together with their high levels of confidence, this suggests that the disagreement described in the previous section is not merely the result of uncertainty, nor confusion regarding the measurement tool (which was the same for all groups), but instead reflects genuine disagreement as to the quality of democracy in the cases in question. It is also consistent with Synodinos’ (1986) finding that individuals’ confidence is not necessarily a reliable indicator of the accuracy of their own retrospective assessments.
Implications and recommendations
What does this mean for our ability to evaluate democracies that have experienced a crisis? Our findings lead us to several recommendations for designers of expert surveys. In general, the V-Dem data are exemplary: the systematic design of questions and response tools enhances standardization; the availability of disaggregated responses and coder confidence levels aids interpretation; and the use of a Bayesian IRT measurement model moderates the effects of coder disagreement that arise from different interpretations of subjective concepts. Other measures of democracy based on expert survey data would do well to follow this lead.
That said, there are still areas in which V-Dem can improve. In particular, we recommend more consistent reporting of coders’ individual traits, including their region of origin/residence, gender, age, partisanship, and familiarity with the case they are coding. Although anonymity is important to ensure honest responses and protect against reprisal, without information on the political leanings of coders, it is difficult to be fully cognizant of their biases. At present, two-third of experts in our samples have not competed the post-survey questionnaire, which also does not contain any questions specifically related to partisanship – a major oversight given the findings of Martinez i Coma and Van Ham (2015: 315–316), whose study of the expert responses used to create the Perceptions of Electoral Integrity dataset suggest that ideological partisanship has a significant effect on experts’ evaluations.
When it comes to aggregating coders’ judgements, V-Dem determines the relative weight of coders’ scores using a Bayesian IRT measurement model, which incorporates data about coders’ confidence and the degree to which coders’ assessments are consistent with those of their peers. In general, this is an effective way to manage problems of disagreement that arise from different interpretations of subjective concepts and keywords (the DIF problem). But when disagreement arises from the fact that similarly qualified and confident experts have substantially different understandings of the same context, this represents a more insidious problem.
Our findings point to two potential areas of concern. First, although V-Dem’s IRT model incorporates judgements about coders’ confidence in their assessments, the subconscious nature of hindsight bias has been shown to lead coders to overstate their confidence, making self-reported confidence unreliable.
Second, the absence of a consistent measure of partisanship/ideological bias among individual coders means that the IRT model has no systematic way to account for this apart from looking at the degree to which a respondent’s assessments are consistent with other evaluations of the same case/phenomenon/year. Given the fact that each indicator is based on the evaluations of a relatively small pool of coders, 18 this leaves open the (currently untestable) possibility that the coders chosen by V-Dem do not represent the full diversity of opinion. This represents a major limitation that is particularly relevant in the evaluation of ideologically polarized pre-crisis contexts.
Hence, while V-Dem’s IRT measurement model is likely to be effective in limiting the distorting effects that arise from coders’ different understandings of measurement tools, it is less likely to offer an effective means of controlling for possible hindsight bias. Thus, at a minimum, we suggest that the post-survey questionnaire should be expanded to incorporate a greater range of questions – especially related to coders’ political leanings – and applied more uniformly, thereby ensuring that a greater proportion of coders actually complete it.
Does this mean that V-Dem data are less reliable than other democratic indices? Far from it. Wherever expert coders are used, cognitive bias is likely to remain a persistent problem. As Schedler (2012: 25) explains, ‘[t]he methodological ideal of impersonal, rule-based, non-judgmental measurement’ ignores both the fundamentally subjective nature of expert judgement, as well as the added discretionary value that experts bring to their analysis – which is the very reason they were consulted in the first place. This suggests not only that unbiased estimation may not be possible, but it may not be desirable either.
Perhaps, because of the difficulty associated with measuring democracies in crisis, pre-crisis periods in other indices are routinely coded as missing in other large datasets. 19 However, rather than overcoming hindsight bias, treating pre-crisis periods as missing data represents a serious problem for those investigating democratic decline, since it is precisely this period that is vital to understanding what precipitated the crisis, identifying potential warning signs of future crises, and gauging the extent of both the crisis and any subsequent recovery. Despite the problems discussed here, the fact that the V-Dem data can be disaggregated is a major advantage for end-users. It is also likely that other, retrospectively coded indices are similarly vulnerable to problems of hindsight bias. However, because they do not offer the same glimpse of individual coders’ assessments – nor even a general error term in some cases – the extent of this problem is not as clear.
Conclusion
The utility and presumed objectivity of expert judgements underpins much of academia, from funding and tenure decisions to the peer-review process. Indices based on expert surveys offer an economical, authoritative, and standardized tool that is particularly useful for facilitating comparison across time and space.
Yet, like lay respondents, experts remain vulnerable to both random and systematic errors in judgement. This study examined the effects of hindsight bias when it comes to understanding how democratic crises affect expert evaluations of democratic attainment. Using data from the Varieties of Democracy Project, it compared coders’ assessments of two groups of ‘democracies’ (broadly defined): one in which a crisis occurred, and one where no crisis was observed. It showed that disagreement was significantly higher in the pre-crisis cases than in the corresponding control groups. In spite of this, pre-crisis coders remain confident in their assessments, reporting a similar level of confidence in their scores as coders of the control cases.
In light of the fact that several large datasets, including V-Dem, are currently in the process of deepening and expanding their historical coverage as far back as the 18th century, the problems associated with the retrospective evaluation of democratic attainment that this study has identified are likely to persist or even deepen.
At the same time, however, given the fact that datasets such as V-Dem are updated on a regular basis, it will soon be possible to examine ‘real time’ disagreement among coders of democracies that are currently on the verge of crises which have not yet come to pass. Will coder disagreement in these contexts follow a similar pattern? This remains to be seen.
Finally, we do not examine levels of coder (dis)agreement that occur when authoritarian regimes become democratic. Do experts also disagree when it comes to coding sudden, positive changes in democratic attainment? If so, is this due to coder uncertainty or some other phenomenon? We leave these questions to future studies.
For the present, this study suggests that scholars who use expert survey data to study processes of democratization and democratic backsliding should exercise caution in doing so. Our findings not only suggest that coders’ knowledge of a crisis affects the way in which they evaluate the events that proceeded it, but also that coders may not be fully aware that this is happening – something that end-users should be aware of, especially if the quality of democracy in the pre-crisis period is a central focus of their argument. In the case of democratic crises, this poses a challenge for those seeking to understand what precipitated the crisis, to identify potential warning signs of future crises, and to gauge the extent of both the crisis itself and any subsequent democratic recovery.
Supplemental Material
POL914571_Appendix – Supplemental material for Hindsight bias in expert surveys: How democratic crises influence retrospective evaluations
Supplemental material, POL914571_Appendix for Hindsight bias in expert surveys: How democratic crises influence retrospective evaluations by Laura Levick and Mauricio Olavarria-Gambi in Politics
Footnotes
Acknowledgements
The authors thank the Post-Doctorate Program of the Universidad de Santiago de Chile’s Vice-Rectory of Research, Development and Innovation for supporting this research.
Universidad de Santiago de Chile, Usach. Agradecimientos Proyecto POSTDOC_DICYT, Código 031752OG, Vicerrectoría de Investigación, Desarrollo e Innovación.
Thanks also to Carsten-Andreas Schulz for his detailed feedback.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research received financial support from Chile’s National Commission for Scientific and Technological Research (CONICYT) through Fondecyt grant No. 1160626.
Supplementary Information
Additional supplementary information may be found with the online version of this article.
Appendix B1: Democratic crises.
Appendix C1: Democratic control groups (sub-sample).
Appendix D1: V-Dem electoral democracy index components.
Appendix E1: Percentage difference in variance.
Notes
Author biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
