Abstract
Picking up on Olof Hallonsten’s contention that contemporary science evaluation is ‘mostly counterproductive’, we argue that the contemporary focus on evaluation is antagonistic to innovation or novelty in science, even though innovation is one of the values that evaluation is often supposed to support. In arguing for the antagonistic relation between evaluation and innovation, we consider arguments from the nature of audit and the situational logic of scientific practice.
The language of ‘evaluation’ that is now commonplace in relation to contemporary institutionalized science carries a positive evaluation of its own that makes it difficult to contest. Who would not endorse the importance of evaluating the value and significance of scientific work? – especially when decisions must be made about the allocation of resources to support it.
Yet although the language of ‘evaluation’ (much like the associated language of ‘evidence’) derives its rhetorical power from the broad set of meanings attached to the term, its actual employment in contemporary scientific contexts is typically tied to a much narrower and more specific sense: to evaluation as a particular sort of institutionalized activity undertaken with respect to particular kinds of ‘value’ and assessments of value. It is a sense of ‘evaluation’ that has arisen, not from the research practices of scientists themselves, but rather from the new managerialist and audit culture that has come to dominate corporations and public institutions, including governments, around the world over the last forty years or so (see Power, 1997; also Shore and Wright, 2015; Travers, 2007). In this respect the practice of science evaluation that is now commonplace is merely part of a larger system of governance and administration that is widespread and is significantly altering processes of decision-making, habits of work, and forms of organization in almost every area of institutionalized life and activity – not only in the laboratory, library, and lecture room, but in the office, the factory, the schoolroom, the law court, the doctor’s surgery, the boardroom, and even the supermarket.
The broader picture is important since it indicates the difficulty of mounting an argument against current science evaluation practices taken on their own. These practices have not arisen from within contemporary science or from within the institutions in which science is embedded, even though they have been adopted by those institutions for their own internal organizational purposes. Moroever, their rise has been driven by ideological and political factors (including the character of contemporary technological capitalism – see Malpas, 2018) more so than anything that has strictly to do with considerations of the efficacy or practicality of scientific work.
Against this background, demonstrations of the counterproductive or contradictory nature of science evaluation practices, or of the extent to which they are inconsistent with the nature of scientific work, are unlikely to have any direct effect in bringing a halt to or significantly altering those practices. This does not mean, however, that Olof Hallonsten’s recent call to stop evaluating science (Hallonsten, 2021), and other calls like it, are futile and without point. Like Canute, we cannot command the tide, and yet that does not mean that we should keep silent about the fact that the tide has been rolling in nor that we should do nothing about it. Moreover, there is also an important question here about the evaluation, in the broad sense of the term, of science evaluation – a scientific question, no less – and it is this question that Hallonsten’s provocation can itself be seen partly to be addressing.
In this spirit, we want to pick up on Hallonsten’s contention that contemporary science evaluation is ‘mostly counterproductive’ (Hallonsten, 2021: 9), and ask, more specifically, whether contemporary evaluation metrics focussed on citation might actually be antagonistic to the emergence of innovation or novelty that is so frequently cited, often alongside the emphasis on evaluation, as one of the key values that contemporary science policy and decision-making ideally aims to promote. There are, in fact, several empirical studies that suggest that evaluation is antagonistic in this way, and that, consequently, evaluation and innovation are in tension with one another.
Before proceeding any further, however, there is an important if perhaps familiar point that should first be noted and that already suggests a difficulty in using the particular sort of evaluation practices that are currently in vogue as a means of encouraging scientific innovation.
Even when the language of ‘audit’ is not explicitly invoked, the contemporary evaluation of scientific work operates within the broad frame of audit practice. That frame imposes specific constraints on the evaluanda, namely that they be sufficiently well-defined, ‘operationally’, to permit independent corroboration of claims about them. As Michael Power (1997: 95) and others already foresaw some time ago, this will often mean that the evaluanda are auditable surrogates for those values that actually matter in the enterprise being evaluated. Thus, to take one obvious example, science evaluation relies on readily auditable metrics associated with citational practices. And while citations are not free-floating in relation to values that really matter in science, such as rigor, methodological innovation, depth of theoretical insight, predictive power, potential technological applications and the like, neither do citations closely track these intrinsic values. And this failure of tracking matters, in particular because surrogates come to be pursued for their own sake and because other goods (such as tenure, promotion, grants, and the like) will be distributed to scientists on the basis of their performance relative to these surrogates, at least in general. At the very least, auditable mechanisms of science evaluation invite gaming the system in pursuit of individual success and are highly prone to giving rise to distortions in perceptions and behaviour. On this basis, there is already a prima facie reason for exercising caution about the extent to which evaluation practices that operate under the constraints of audit will indeed be conducive to the accurate monitoring of what is valued in scientific work, including innovation, or its encouragement.
Publication and grant success are two key factors in underpinning positive evaluations of individuals within institutional scientific settings, and since publication success feeds into grant success, they are also connected. The empirical evidence suggests, however, that publication is more likely when it connects to already published work and when it deploys already established methods within a (typically sub-) disciplinary field. Certain topics, approaches, or figures thus function as attractors so that scientific work, and especially scientific publication, exhibits a ‘clustering’ effect (see D’Agostino, 2019a, 2019b). It is easier for scientists to do this sort of ‘normal science’ work because – or insofar as – such work relies on ideas and techniques that they are already familiar with and competent in the use of. And it is easier for referees and supervisors to evaluate such work because they too already have the skills and knowledge to recognize competent performance. Research and its assessment are harder when the topics at issue and the methods used are not already familiar and established; such topics and methods stretch capacities on both sides of the evaluation interface.
The situational logic that is associated with the institutional context in which scientific work is most immediately embedded thus has an inevitably normalizing or moderating effect that is indeed antagonistic to innovation or novelty. Put simply, scientific work that is innovative and novel is also risky, and the more influenced are researchers by practices of evaluation at the institutional level, the more inclined they will be to reduce that risk in order to maximise the chances of positive evaluation at that level.
This holds even though the more publications that cluster in this way the less likely any specific paper is to become heavily cited. In the immediate institutional setting any publication or citation is better than no publication or citation, and the risk of trying to publish in areas outside of the norm – away from the cluster – is that one will end up, in the short term, with no publication or no citation (see D’Agostino, 2019a, 2019b). Such clustering occurs even without any emphasis on the quantitative evaluation metrics associated with citation analysis or measures of impact but is also reinforced by such metrics.
The clustering effect that is at issue here would seem to be only marginally, if at all, belied by evidence that suggests a decrease in citation concentration (see Larivière et al., 2009). Such a decrease in concentration, or as it might also be characterized, increase in dispersal of citations, is evidence of a change in the pattern of clustering, not of the importance for scientists of their work being embedded in a recognized research cluster. On our analysis, it is indicative of a form of clustering associated with an increase in sub-disciplinary specialization, which has long been recognized as a consequence of individuals’ tactical responses to their situations. As Richard Whitley put it (1984: 295):
Specialization has been encouraged by the large [post-War] increase in numbers of scientists competing for reputations [. . .] Thus scientists narrowed their topics and foci to avoid direct competition but have been able to claim contributions to intellectual goals [. . .].
This analytic point about the antipathy between innovation and evaluation can be made more pointed by looking to bibliometrically-oriented studies of scientific work that shows a high degree of innovation or novelty. So, for instance, if one assumes innovation or novelty is tied to a degree of conceptual diversity, and assuming a measure of such diversity can be identified in the diversity of the sources on which a publication draws (see Wang et al., 2016), then: (1) more ‘diverse’ publications are less frequently cited than more ‘intellectually compact’ papers in the ten years or so following their appearance, but (2) are more highly cited on average than the compact papers in the longer term, (3) are disproportionately represented among the highly-cited papers in the longer term, and (4) are more highly cited (than the compact papers) by authors who are highly cited. In addition, it appears that (5) the ‘diverse’ papers are less well represented in the highest-prestige journals, notwithstanding points (2), (3), and (4) (see Wang et al., 2016). Once again, given the institutional context that determines the situational logic of scientific work, reward is more likely to flow to those papers that are more highly cited in the short- than the long-term and so to papers that are more ‘compact’ rather than ‘diverse’ (see also Uzzi et al., 2013).
It seems well-established, if not always widely acknowledged, that innovation has been declining worldwide over the last century (see eg. Huebner, 2005). Some have argued that the increasing reliance on quantitative evaluation is one of the factors that has contributed to that decline in recent decades (see Bhattacharya and Packalen, 2020). Yet regardless of whether evaluation practices are directly implicated as a factor, the evidence of innovation decline ought to make us wary of structures and practices that reinforce the inevitable tendencies towards normalization. Setting aside the complications of the larger ideological and political context, the fact that there does indeed seem to be an empirically demonstrable as well as a theoretically articulable antagonism between innovation and evaluation (in its dominant contemporary forms) ought therefore to discourage reliance on such evaluation and lead to the exploration of other ways in which science can more adequately be managed and supported.
Footnotes
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
