Abstract
Evaluation as a free-standing discipline arose out of the ashes of World War II, a time of optimism, when government turned to the academy to guide public policy. The evaluation pioneers shared a bracing vision: a search for truth in the public interest. Seventy years later the glitter has faded, and disenchantment has taken hold. Evaluation, a quintessential public good, has become a market good, and eminent evaluation thinkers are asking the same questions about evaluation that they have been routinely asking of others—with sobering results. Yet, countervailing currents and turbulent streams lie just below the surface. Once a tipping point is reached, a new wave of evaluation diffusion will begin to curl. What might it look like?
“The fate of our times is characterized by rationalization and intellectualization and, above all, the disenchantment of the world (so that) precisely the ultimate and most sublime values have retreated from public life”
Introduction
This opinion article is in five parts. First, it sketches the history of evaluation and describes the self-doubt that the evaluation community is currently experiencing. Second, it relates the roots of this disenchantment to the perennial tension between reason and faith. Third, it uses Weber’s concepts to illuminate the on-going shift in attitudes within the evaluation community. Fourth, it unpacks the definition of evaluation to show that a tighter conception of its remit and a sharper ethical focus would be needed to restore faith in the discipline. Fifth, it speculates about the professional future of the discipline.
From faith to disenchantment
Evaluation as a knowledge occupation emerged at a time of optimism and belief in government. Its creators shared a common view about the mission of the fledgling evaluation discipline: to use the power of reason and observation to enhance the effectiveness of social programs. By connecting the social sciences to policy making, the pioneers strove to inject rationality in the untidy world of politics.
Capturing the spirit of the times, Carol Weiss (1993) called evaluation “a rational enterprise taking place in a political context.” Thus, the ideal “experimenting society” proposed by Donald Campbell (1971) was an open plea to policy makers to learn from experience and subject social programs to systematic evaluation. Toward this end, Campbell developed a rich tool kit. But at a time of rapid social change, Campbell’s technocratic vision of the experimenting society did not live up to the democratic aspirations of evaluation practitioners and users.
Hence, new evaluative processes were fashioned to improve the workings of democracy (House, 2005) and the scientific wave of evaluation practice that had swelled in 1960 was upstaged by a participatory wave of progressive social action that washed over the discipline in the 1970s—the halcyon days of inclusive, dialogue-oriented value-driven evaluation (Vedung, 2010). Both waves were propelled by a faith in evaluation as a force of good in society.
Inevitably, political counter-currents materialized so that this exceptionally innovative period of evaluation history came to an end when a neo-liberal wave engulfed the evaluation discipline in the early 1980s. This is when the New Public Management (NPM) movement introduced market thinking into government and value-free management consultants invaded all sectors of society to serve senior executives, champion “free market” ideas and celebrate free-wheeling capitalism (Parker, 2018).
The NPM ethos has proved resilient: it is sustaining the current evidence-based wave that emerged around 1990 and that still characterizes evaluation diffusion today. Its prime movers advocate evaluation geared to value for money and incontrovertible evidence of verifiable “results,” a notion that harks back to the experimental phase of evaluation history. With randomized control trials back in vogue (as confirmed by the recent Nobel prize in economic sciences), methodological skirmishes, while dormant and notionally put to rest under the mixed methods banner (Stern et al., 2012) may once again flare.
Evaluation today
Far more than the well-established knowledge professions, evaluation has been exceptionally vulnerable to the threats posed by the contemporary operating context. This is paradoxical since, carried by a headlong pursuit of scientific rationality at the service of the public good, the evaluation prime movers had developed a host of theories, models, analytical tools, guiding principles, support networks, training programs, competency frameworks, and specialized publications (Scriven, 1991a).
Furthermore, under the sway of globalization, evaluation has spread its wings across borders and evaluation practice has become genuinely “international in the sense of being at the same time more indigenous, more global and more transnational (Chelimsky and Shadish, 1997).” Why then did self doubt emerge?
Paradoxically, it may be because evaluation has become omnipresent and a victim of its own success: “In recent years, we have witnessed a boom in evaluation . . . It is as if there is no limit to the feedback loops . . . as if the insatiable evaluation monster demands more food every day” (Dahler-Larsen, 2012).
There is little doubt that evaluation has become integrated as an administrative routine in all sectors of the economy: the utilization focus of the evaluation enterprise induced many evaluation practitioners to acquire “skin in the game” (Taleb, 2019). In pursuit of greater influence, they have jumped into bed with decision-makers thus forsaking their autonomy, devaluing the accountability dimension of their work, and blurring the distinction between evaluation, monitoring, and management consultancy.
Co-opted, institutionalized, and routinized, evaluation is now shaped by buyers’ preferences, the range of evaluation questions has become restricted and manager-oriented evaluations no longer privilege the weaker segments of society. In this operating context, evaluators are induced (or seduced) to share their clients’ perspectives, while giving lip service to the public good: they cannot afford to bite the hands that feed them.
This shift in evaluation policy directions has blurred the boundaries between evaluation and other knowledge occupations. Hence, evaluation in the public mind has become conflated with quality assurance, performance assessment, internal and external auditing, inspection, and other means of social control that are widely perceived as costly, ritualistic, and disruptive (Power, 1994).
Intrusive oversight, detailed record keeping, intense bureaucratic scrutiny, constant pressure to demonstrate rapid results, insistence on frequent progress tracking, and mandatory use of simplistic performance measures currently prevail. They provide illusory comfort to managers faced by the realities of an uncertain and turbulent operating environment that precludes easy quantification (Natsios, 2010).
Indicator proliferation has taken hold under the spell of business school doctrines that marginalize expert judgment, hinder local ownership, discourage institution building, and neglect the flexible, long-term associations required to assure sustainability. In parallel, public skepticism regarding government effectiveness has boosted the demand for hard evidence, tangible results, value for money, and so on.
Evaluation, once a public good, is now bought and sold in a market where evaluators share control over their work with commissioners. To be sure, value-laden evaluation models that emphasize social justice, democracy, and inclusivity are legion. But they are not in widespread use. They have been upstaged by a utilization focused culture dedicated to goal achievement.
In parallel, evaluators, pushed toward the periphery of policy making, are increasingly relegated to the fulfillment of managers’ needs for data gathering and interpretation (Eyben, 2013). Furthermore, the cost-effective computer algorithms of the new information economy, while often flawed and biased, now guide decision-making in public and private organizations so that theory free, value-free, data scientists are increasingly displacing evaluators in the competitive marketplace.
Self-doubt takes hold
Given the above trends, it is not surprising that self-doubt should have become rampant or that eminent evaluation thinkers would strike a critical stance. They are now asking the same questions about evaluation that they routinely ask of others. Does evaluation “work”? What has been achieved and at what cost? Whatever happened to evaluation? (Furubo and Stame, 2019: xvi).
The results of their self-examination are sobering. According to Thomas Schwandt (2019), evaluation has yet to define and embrace an explicit understanding of its professional ethos and that “despite a repertoire of shared knowledge there remains significant disagreement in the field about what it means to practice evaluation.”
Adopting an equally skeptical position, Kim Forss views evaluation as a “systems-preserving activity . . . an intellectual effort which is inherently conservative and that assists in defending rather than challenging the powers that be, the established wisdoms, the current technologies and administrative practices” (Forss, 2019).
As for Peter Dahler-Larsen (2012), he deplores the high transaction costs and the unintended effects of linear evaluative thinking on creativity and innovation and concludes that “some of the self-congratulatory rhetoric of the evaluation industry may be unwarranted. It is time to consider . . . whether the marginal utility of evaluation may be decreasing and whether there are sometimes good reasons for evaluation fatigue.”
In this critical climate of expert opinion, the conception of evaluation as a pragmatic, collaborative enterprise jointly owned by evaluators, decision makers, evaluation commissioners, and so on has gained ground and the hallowed principle of evaluation independence has taken a back seat. In pursuit of greater utilization, evaluation practice has become adaptable, flexible, and opportunistic.
Evaluation is now conceived as an enterprise rather than a vocation. Shaped by the interplay between commissioners and suppliers of evaluation services, evaluation has become a business venture rather than a special calling—a specialized form of action research or a managerial consultancy service: “one tool among many for the improvement of policies, learning and social change” (Furubo and Stame, 2019: xv).
What then are the prospects of the evaluation occupation? Will it be able to self-generate a broad-based consensus about its identity, logic, and philosophy of operation? Will it survive as a distinct knowledge occupation given the cooptation of evaluation ideas by the auditing, accreditation and management consultancy disciplines? Should it attempt to achieve the status of a self-managed profession? These existential issues have come center stage on the evaluation scene. As a modest contribution to the suddenly thriving sociological analysis of evaluation, this article seeks to identify the underlying sources of the contemporary evaluation malaise.
Enlightenment dialectics
The interface between faith and reason has always been fraught with ambiguity and tension. While faith and reason co-existed at the creation of the evaluation discipline, tensions arose once neo-liberal economics captured the commanding heights of public policy. The advent of evaluation as a distinct occupation evokes the Enlightenment when the social sciences emerged with the heady promise of putting rationality at the service of the public good.
But raw power and ideology always intervene so that the “Century of Light” degenerated into the Terror of the French Revolution. According to Justin E. H. Smith (2019), reason has a dark side: the harder the struggle for reason, the more likely the emergence of unreason. The celebration of savage capitalism, the extrapolation of Darwin’s evolution theory to society (“survival of the fittest”) and the persistent lure of racial politics have been traced to the decline of religion and the consequent rationalization of modern society (Bauman, 1989).
On the one hand, prominent Enlightenment philosophers—for example, Voltaire, Rousseau, and Diderot— attacked and ridiculed religion. 1 On the other hand, René Descartes believed that one can prove the existence of God through logic and Blaise Pascal famously declared that the heart has its reasons which reason knows not. Straddling the extremes, David Hume (1739) acknowledged that being reasonable means accepting the limits of reason. Thus, the Enlightenment was a “stage on which fierce and unpredictable battles were fought between reason and passion or, later, among the various passions” (Hirschman, 1981).
Then as now, no consensus emerged about the social impact of individual profit seeking behavior (Hirschman, 1977). Thus, Adam Smith (1776) asserted that the social good is enhanced by the pursuit of profit even though every individual neither intends to promote the public interest, nor know how he is promoting it: he is in this, as in many other cases, “led by an invisible hand to promote an end which was no part of his intention.” Conversely, Adam Ferguson (1767), a Scottish Enlightenment thinker, deplored the spirit of commercial societies, where the bands of affection are broken and where corruption, luxury and desire for tranquility favor despotism.
The tug of war between proponents and detractors of the Enlightenment and of free market ideologies continues to this day. Popular social scientists such as Steven Pinker (2018) and academic statisticians such as the Rosling et al. (2018) have mobilized reams of data to demonstrate that Enlightenment ideas and liberal worldviews grounded in reason have been hugely beneficial and will continue to generate major improvements in human wellbeing. On the other hand, such eminent philosophers as John Gray (2019) attribute nihilism, colonialism, and twentieth century totalitarianism to the excesses of heedless rationalism and faithless modernity.
Evaluation has not been spared the widespread discontent associated with the rise of modernity, the decline of traditional values and the attendant fraying of social ties. Vulnerable to manipulation by vested interests, evaluation has been elbowed out of its market by other disciplines. Fashioned into an instrument of social control, it has been assailed by liberal critics. Captured by micro-economists who equate evaluation with experimental methods, it has become indistinguishable from evangelical scientism.
The anatomy of disenchantment
In evaluation as in other social pursuits, the mental processes that facilitate ethically inappropriate behaviors are similar to those that oppose individuals’ instinctual search for gratification and the wholesome imperatives of societal life in a modern civilization. They trigger guilt and even neurosis, according to Sigmund Freud (2014 [1930]). Thus, the gap between the lofty aspirations prevalent at the creation of the evaluation discipline and current market realities underlies the discontent that plagues much of the evaluation community today. 2
Why reach out to Max Weber to assess the evolving force field that has reshaped evaluation and contributed to a pervasive crisis of confidence among its practitioners? Because this eminent German philosopher, jurist and political economist (1864–1960) is a major contributor to social theory. Arguably, he laid the intellectual groundwork for the rise of the evaluation discipline since he advocated the study of social action through interpretive and empirical means and gave pride of place to the purpose, meaning and values that individuals attach to their own actions.
At the outset of his most famous essay, Weber (1958a) wondered how money making and commercial pursuits that had previously been despised as manifestations of greed and avarice became honorable in the modern age. He discovered that the rise of the Protestant Ethic was the result of complex interactions between theoretical and practical rationalities (Carroll, 2011).
Contemplation was demystified, action became privileged and the sacramental mediation of salvation was rejected. This paved the way for modern capitalism. Work became valued for its own sake. Triggered by disenchantment of the world in the wake of scientific advances, a decisive shift of religious focus toward rational action gradually transformed everyday life and laid the foundations for the triumph of capitalism.
First, the scientific method induced more precise and abstract concepts, thus rationalizing away magical images of the world. Next, practical mastery of the world was facilitated by means-ends mental models. This is when believers in predestination began to look to business success as a sign of divine favor while practicing restraint and striving to comply with strict morals in their private lives.
The path to salvation shifted from a contemplative “flight from the world” toward an active ascetic “work in the world” ideal. As a result, the acquisitive urge that fueled the growth of commercialized society was contained by social norms that eschewed ostentation, encouraged charity, imposed bounds on conspicuous consumption and promoted investment that would generate economic opportunities for the working class. Hence, the popular acceptance of a new hierarchy of wealth.
But with the rapid decline of religious faith in modern society, morality was “outsourced” through reliance on laws and rules, a development that in turn has induced “gaming” to get around the regulatory system. In parallel, the emergence of flexible labor markets and of the gig economy have turned employees into interchangeable agents, reduced social trust, and coarsened the fabric of society (Milanovic, 2019).
Abstract theories, concepts, and norms have become increasingly influential and the obsessive pursuit of managerial goals shaped by vested interests has yielded moral insensitivity and undermined respect for human dignity. Here as elsewhere, strict adherence to norms and standards, have ended up sacrificing human values at the altar of predictability and societal efficiency. As secularization eroded moral values in the public sphere and the need for social order, efficiency, and predictability yielded proliferating bureaucracies, reliance on value-free social doctrines and neglect of ethical standards.
Thus, Weber’s sociological diagnosis of modernity helps unpack rationalization processes. For example, tolerance of unethical behavior under the guise of social imperatives is characteristic of formal, bureaucratic forms of organization that, in their extreme incarnations, become dictatorial, inhumane and ultimately irrational. This is why Weber stressed that only rationality that privileges moral action has the potential of subjugating the risks to society in an era disenchanted from faith. But he was not optimistic about its prospects given the relentless rise of soulless bureaucracy and value-free politics.
These observations still resonate and they provide a useful backdrop for an examination of the contemporary predicament of the evaluation enterprise. Massive increases in inequality and the existential risks associated with environmentally unsustainable economic policies have added fuel to the fire of popular anger. The future is already here. According to Ernest House (2016), “conflict of interest in evaluation has increased in several fields in which evaluation plays an important role, including pharmaceutical evaluation, social and education evaluation, and financial evaluation.” It is time for evaluation to reconsider its place in the world.
Back to basics?
According to Michael Scriven (2007), evaluation is “the process of determining merit, worth,
The policy content of any evaluation is shaped by the weight it ascribes to each of the three evaluation dimensions. First, the merit dimension of the evaluation definition evokes Weber’s concept of value-free formal rationality. Evaluators consider that an evaluand displays merit when it complies with pre-determined policies, norms and standards. Thus, merit is intrinsic to the intervention being evaluated. It is about doing things right.
For example, product evaluation consists in (1) establishing criteria of merit that privilege pertinent evaluand dimensions; (2) constructing performance standards that reflect these criteria; (3) measuring performance against these standards; and (iv) synthesizing the results into a judgment of merit (Scriven, 1981).
Second, and by contrast, the worth criterion is extrinsic. Specifically, assessing worth involves checking whether or the intervention is doing the right things for individuals identified as stakeholders. It addresses the extent to which a social action meets stakeholders’ aspirations and needs without necessarily complying with predetermined merit standards. Thus, worth evaluation reflects Weber’s concept of value-free practical rationality since it has to do with satisfying the requirements of the self-interested individuals and groups framed by the evaluation.
Drawing unambiguous summative conclusions from a combination of worth-oriented and merit-oriented evaluations is not a trivial task since the needs and preferences of individuals and groups affected by a social action are bound to differ. The boundaries set around the evaluation and the values mobilized to assess the evaluand play a large role in evaluative judgments. There is no easy escape from collective action dilemmas, for example, Kenneth Arrow’s impossibility theorem which states that if the preferences of two stakeholders or more need to be satisfied when choosing among three options or more then it is impossible to select goals that satisfy all stakeholders (Maskin and Sen, 2014).
This is why significance is the end-game when all pertinent data and evaluative judgments, including the size, importance and transformative effects of the social action, are synthesized to reach an overall judgment of value (Scriven, 1991a). Significance assessment identifies the nature and weighs the extent of the gaps between merit and worth and strikes a judicious balance that seeks to value the social importance and transformative impact of the evaluand. Significance is therefore about doing good as well as doing right from a public interest perspective, with special consideration for the interests of the most disadvantaged groups in society.
As a result, there are many ways to assess significance depending on the perspectives of the individuals or groups involved in the social action as well as those of the evaluator. Only ethics can address this conundrum. Thus, significance embodies but also transcends merit and worth assessment. According to Deborah Fournier (Mathison, 2005), conclusions made in evaluations encompass both an empirical aspect (that something is the case) and a normative aspect (judgment about the value of something). It is the value feature that distinguishes evaluation from other types of inquiry, such as basic research, clinical epidemiology, investigative journalism, or public polling.
Revealingly, a tighter evaluation definition that requires evaluators to address all three dimensions of their craft was also put forward by Scriven (1991b): “Evaluation is the process of determining the merit, worth,
The next wave of evaluation diffusion
Evaluation has acquired all the characteristics of a distinct discipline in its own right. It also enjoys trans-disciplinary features that allow it to support all the social sciences through a vast repertoire of well-honed approaches and processes. But the world has changed. We are about to enter a new evaluation age. How will the evaluation discipline adapt to a volatile, interconnected, market-driven operating environment?
Is evaluation equipped to respond to emerging policy priorities? Will it rise to the challenges of unprecedented and growing inequality? In a world where social impact has become a major preoccupation of public, private, and civil society actors, will the changed operating context drive new policy directions and elicit a transformational agenda for the evaluation community?
In recent decades, powerful political forces have favored value-free evaluation. As a result, social policy has neglected inclusion, equity and the welfare of future generations. The single narrative of free market ideology exerts a major influence on society. But rising public discontent with the inequitable and unsustainable outcomes of the current policy mix has generated a pent-up need for value-driven, independent, no holds-barred, transformational and ethically valid evaluations.
Nurturing and meeting this latent demand holds the key to the renaissance of evaluation as a value-driven discipline. Given the challenges associated with social inequality, climate change, and totalitarian movements, the world would be well served by a greater supply of value-committed, independent, ethical evaluation services grounded in tolerance and dialogue. In particular, and beyond existing approaches and tools for evaluating projects, programs, national policies, and plans, global problems are in dire need of global systems evaluation as proposed by Michael Patton (2019). This will require evaluators’ fulsome commitment to tighter evaluation ethical standards.
Evaluation ethics
The ethical guidelines issued by evaluation associations are useful but narrow: they concentrate on individual evaluators and neglect the evaluand. According to Nicoletta Stame (2018): “Too often evaluation confines itself to assessing whether objectives set by program designers are achieved.”
In Michael Scriven’s (2016) words, there is [. . .] a need for a considerable expansion of the present norm [. . .] We must do a specifically ethical analysis of each evaluand in the relevant contexts, because [. . .] doing that is part of our professional responsibility, as is the requirement that we give scientifically acceptable reasons for our stance, including its ethical assumptions.
Exploring this last frontier of evaluation practice implies tighter boundaries around evaluative inquiry. They would be more demanding by no longer allowing exclusion of basic ethical concerns. This would imply a more ambitious remit regarding the ethical assessment of evaluands. All evaluations would make adequate room for expert estimation of the indirect and unintended social and environmental effects and strengthened evaluation guidelines would address the ethics of all evaluation participants. 3
This means that evaluators would be enjoined to refuse evaluation assignments intended as subterfuge—for example, evaluations commissioned to delay needed action, to duck responsibility, for window dressing or for public relations. Evaluators would take full responsibility for their work and subject evaluation terms of reference to critical review so that the evaluations they agree to carry out do not strengthen the hold of powerful elites while ignoring or neglecting groups with legitimate claims (Schwandt and Gates, 2016).
The imperative of professionalization
But is this more than a dreamy aspiration? Can it be reconciled with the fragile status of evaluators facing stiff competition in a market place dominated by the monopsonist power of evaluation commissioners? Can evaluation live up to the dictionary definition of a profession: a vocation, a calling requiring advanced knowledge or training in some branch of learning or science?
Thomas Schwandt (2018) has rightly emphasized that too little attention has been given to the normative characteristics that are unique to evaluation. He has powerfully argued for a commitment to democratic professionalism. But professionalism is a heroic challenge without professionalization. For instrumentally rational, ethical evaluation to prevail it must exert influence: the evaluation occupation will have to upgrade its social status by moving upward on the totem pole of occupational groups.
To do so brand differentiation is critical, and the sociology of professions has established that the status of any expert occupation is closely connected to its autonomy and capacity for self-management. Tightening evaluation guidelines and principles will not be enough. Thus, looking at the evaluation enterprise through an economist’s lens, (Nielsen et al., 2019) have argued cogently that the growing public demand for accountability has had fundamental implications for the evaluation occupation.
They have observed that the evaluation enterprise as currently organized does not control the supply of evaluation services through designation, credentialing, or certification so that it has been unable to protect its brand in a lopsided market dominated by buyer power. As a result, the evaluation market has been invaded by other knowledge occupations (policy research, management consultancy, performance auditing and now data science) and the boundaries between evaluation and other knowledge occupations have become porous.
In order to escape this predicament, evaluation faces a collective action dilemma that can only be addressed through purposeful collective action, that is, professionalization. According to the sociology of professions, market control, remuneration and social status are positively correlated. Professionalization translates “one order of scarce resources—special knowledge and skills—into another—social and economic rewards” (Larson, 1977).
The lessons of history are unambiguous: professionalization cannot take place unless the possessors of a specialized body of abstract knowledge form themselves into a cohesive group with the authority to limit access to its ranks (through competency tests), control the supply of its services (through legitimate norms and standards), and evoke disinterestedness and public service ideals in ways that promote the respectability and social standing of its members (Cooper et al., 1988).
Indeed, scholars define professionalism as a set of institutions which resist the tyranny of the market and the excesses of state power. It allows members of an occupational group to make a good living while controlling their own work (Freidson, 2001). Only professionalized, independent evaluation would be in a position to upstage fee-dependent and manager-driven evaluations. In parallel, social utility rather than utilization would become the overarching benchmark of high-quality evaluation.
This means that evaluation disenchantment will not lift and that evaluation will not be perceived as a vocation unless evaluators opt to control access to their ranks, empower themselves and take full charge of their work. Once they do, it is likely that the currently permissive definition of evaluation will be revisited; that the ethics of evaluands will be addressed by upgraded evaluation guidelines; and that the value-driven evaluation discipline will acquire the occupational autonomy it needs to maintain the integrity of its processes and to avoid its capture by vested interests. There is no shortcut.
Footnotes
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
