Abstract
The text comprises two main parts. The first part reflects on the status of simulation as scientific instrument and clarifies the relationship between modeling and instrumentation. The second part aims at investigating how simulation is organized. My main claim is that a turning point for simulation occurred around 1990 that signals a new quality of simulation in terms of its social and cognitive organization. Computational chemistry will serve as primary example. So-called density functional theory (DFT) has not yet received much attention from the side of science studies, whereas the 1990s turn propelled it into the arguably most widely used theory in all of chemistry and physics. This success, I argue, is based on networked and cheaply accessible computers as well as on how DFT is socially and cognitively organized, much like the power of ants depends on their organization.
Introduction
Elephants and ants are very different animals. The elephant is the largest animal on land, it is hard to overlook and earthshaking when it moves. Other species escape attention much more easily; and here I think in particular of ants. Each single ant is a tiny creature, but they are very busy beings and build large colonies. For sure, elephants rarely are lonesome rangers, but organize themselves in herds, i.e. assemblies of individual members. Ant colonies, however, exemplify a different type of social organization where specialized components work together and function in a way totally different from what a single ant could achieve. Social organization matters to a remarkable extent. In a competition of weight, no elephant would consider it worth competing with an ant. However, the picture changes when accumulating across individuals. If all the elephants and ants in the rain forest were to meet and hold a competition to see which of them is weightier, the giants would lose outright. In fact, the average cumulated weight of ants per square kilometer exceeds that of elephants by a large factor!
How does this little story about rain forest statistics feed into an article on scientific instrumentation? Well, looking at species of animals is not too different from looking at researchers in science. Despite all obvious differences, there is one important parallel and this will guide the following investigation of how instrumentation and social organization affect scientific activity. Historians, philosophers and other students of scientific practices tend to concentrate on elephants, like landmark theories, and are in danger of missing the ants, i.e. practices that are based on less prominent theories and models, but outweigh their more prominent colleagues.
In the first section below, I will argue about computer simulation, one of the most important scientific instruments of our times. Of course, the computer is surely an elephant – visibly (and invisibly) pervading virtually all areas of science and even the broader culture. What characterizes this device? Often it is seen as a means to speed up calculation. And indeed, the computer is a number-crunching animal of gigantic appetite. Accordingly, scientists can feed more and more into it and get out ever more voluminous results, better approximations and so on. Science studies get trapped if they take this as the whole picture. The following investigation will take on a different perspective. It will scrutinize computer simulation as an instrument that socially and cognitively organizes and structures scientific practices. Ignoring this organization would be like ignoring the ants.
The text comprises two main parts. The first part (the second section below) reflects on the status of simulation as scientific instrument and clarifies the relationship between modeling and instrumentation. Basically, simulation helps modeling to fulfill the essential instrument functions of detection, metrology and control (Marcovich and Shinn, this issue). The versatility of simulation modeling is identified as a pivotal factor. This versatility is tied to simulation methodology and makes essential use of a feedback loop in modeling that allows model behavior to be adapted and also creates a characteristic tension between imitation and representation.
The second part of the article aims to investigate how simulation is organized. This undertaking requires a two-tiered strategy. The third section below deals with tier 1 and presents data that indicate two turning points in the history of simulation. The first turn is what everybody expects, namely that simulation methods started to flourish when digital computers became available in the 1960s: it signals the birth of an elephant. There is a second turn of equal dimension, however, occurring around the 1990s. What happened at this turning point? My main claim is that the 1990s turn signals a new quality of simulation in terms of its social and cognitive organization. In the metaphorical language of the jungle used above, it is the turn from elephants to ants.
Substantiating and supporting this claim is the task of tier 2 (the fourth section below). Giving a substantial argument requires zooming into a typical case and explaining how simulation methodology and organization are intertwined. My choice is computational chemistry, and, in particular, so-called density functional theory (DFT). It has not yet received much attention from the side of science studies, whereas the 1990s turn propelled it into the arguably most widely used theory in all of chemistry and physics. This success, I argue, is based on networked and cheaply accessible computers as well as on how DFT is socially and cognitively organized, much like the power of ants depends on their organization. The concluding section will identify some challenges posed by these findings.
Computer simulation as instrument
Is computer simulation an instrument? It is not obvious this question should receive a positive answer. As the name says, computer simulation is a kind of simulation generated with the help of the computer. The computer itself is a device for performing calculations or, more precisely, for executing formal algorithms. Simulation is different from mere calculation; the centerpiece of simulation is rather simulation modeling. It features special and characteristic ways of building models and working with them. The question hence might be refined: Is simulation modeling an instrument?
My answer will be yes, but there is a conceptual difficulty that has to be addressed first. Modeling has been there for a long time. Among the family members of modeling are not only material models, but also intellectual models, like mathematical models. Simulation models belong to the latter sub-family of intellectual models. What instrument function does modeling comprise? Has this function been around since intellectual or mathematical models were in use? Or is there a special function of simulation modeling where simulation is different from other kinds of modeling?
A good reference point for dealing with these questions is Anne Marcovich’s and Terry Shinn’s contribution to this issue. There they discern three categories of instrument function, namely detection, measurement and control (Marcovich and Shinn, this issue). Modeling does not add a fourth category. The point rather is that simulation facilitates and boosts modeling to a level where it can fulfill these instrument functions.
A pertinent example that illustrates the point is atomic research. Designing and maintaining the military arsenal required nuclear testing. The USA promoted a test ban after they were confident they could replace the physical explosions and the usual material instruments by simulation methods that would detect, measure and control the (now virtual) explosions to a satisfying degree. Nowadays, scientists and engineers in many different areas routinely use simulations to execute the three instrument functions. In other words, simulation turns modeling into an instrument and hence can aptly be called an ‘instrument of instruments’. 1
A crucial component of why simulation models can function as instruments consists in how adaptive these models are. In the remainder of this section, I take a closer look at two facets of this feature. Conceptually, there is a tension between representation and imitation typical for simulation models. Methodologically, a feedback loop in modeling facilitates the great versatility.
Representation vs. imitation
I would like to point out that a particular tension between imitation and representation characterizes simulation in a quite general sense. Küppers et al. (2006) elaborated on this tension as structural versus functional accounts:
Alan Turing (1912–1954) situated the validation of a simulation in the setting of an ‘imitation game’ (Turing 1950). An interrogator had to pose questions via telex and received responses from a human and a computer. Could he discriminate which set of answers was given by the machine and which by the human? A computer whose answers are indiscernible would imitate the answers of a human being and thus pass the ‘Turing Test.’ To date, no computer has succeeded! Hence the imitation game that constitutes the Turing Test afforded a sophisticated arrangement to filter the functional equivalence out of the plenitude of theoretical possibilities. A computer has to behave like a human to pass the test, but the potential success has to be independent of the mechanisms that are employed to simulate. Turing, therefore, adhered strongly to a functionalist perspective. … This standpoint has been challenged by a position that can be called ‘structural’ and that maintains that a valid simulation has to be structurally valid, that is, has to be based on a simulation model that represents [emphasis added] the structure of the system under investigation. This problem has been taken up in J. Searle’s (1980) famous and controversial ‘Chinese room’ argument that boils down to the point that imitation [emphasis added] of intelligent behavior is not itself intelligent behavior. … Should [a test for artificial intelligence] provide a functional or structural imitation of cognition? (2006: 14–15)
This lengthy quote describes an instance with general significance. Where simulation is used as an instrument, the tension between representation and imitation (or between a structural and a functional account) is present. If you remind yourself of Terry Shinn’s characterization of research instruments, like genericity, and disembedding – re-embedding, it does not come as a surprise that the instrument does not work on the basis of representation alone. The versatility of imitating several dynamics of interest, contrary to representing one particular dynamic, correlates with the facility of being embedded in new contexts, contrary to being tied to one particular context.
Feedback loop in modeling
A look at the methodology of simulation modeling shows why imitating phenomena is a persistent principle of simulation that sustains the tension between imitation and representation.
There are certain classes of simulation models, like artificial neural networks, which aim at imitation from the outset. Their point is exactly to connect input and output in a highly implicit way that is based on extensive parameter adjustments – ‘learning algorithms’ – instead of explicit mathematical equations. If such models are successful in simulating a phenomenon, it is because their behavior can be adapted to the desired input–output patterns, not because the models would represent the essential structure of the phenomena of interest. The role of imitation therefore is straightforward. The more difficult case for my argument is when the researchers start from a theory in the form of mathematical equations, like thermodynamics for example. 2 The role of adaptation should be smaller since the models are based on a strong theoretical structure. In the remainder of this subsection, I want to concentrate on this case where a certain quantity of interest, say x, is represented in a theoretical model by xmod.
Models of thermodynamics, to keep this example, are older than the computer. So what is the point of simulation? With simulation methods, researchers can build and handle complex models in a flexible way. At the same time, the modeling process becomes broader and more sprawling. Addressing problems by simulation connects three important issues: setting up the theoretical model (suitably based on the theory of thermodynamics), implementing and executing it on computers and analyzing the results. 3 Obviously, the discrete modeling and implementation steps are particular to simulation. They transform the theoretical model with its continuous equations into that sort of discrete object computers can handle. This transformation includes specifying algorithms for solving the (now discrete, i.e. step-wise) equations, coding in a particular programming language, and compiling on a given constellation of soft- and hardware.
The analysis and evaluation of the theoretical model via comparison to the target system is still the standard rationale. In general, the quality of a model depends on two aspects that counteract each other. It depends both on adequacy of representation, else the model would not yield results revealing anything about the target system, and tractability, which is prerequisite for obtaining any result at all. Here is where computers have changed the picture. They can handle very long and convoluted iterative algorithms that would be intractable for human beings and, hence, make models tractable which otherwise would be useless.
Although the modeling process is enlarged, it makes complexity more manageable than it has been on the basis of theoretical (formal, mathematical) models alone. The key is a feedback loop that renders models adaptable to an extraordinary degree. Making use of adaptability, in turn, affects how simulation is socially organized. First let me explain how the feedback loop operates. Consider researchers are interested in a quantity xreal of some target system. The quantity of interest could be the elasticity of a mixture of certain substances. Chemical engineers will use thermodynamics theory to specify a model based on so-called equations of state that determine the desired quantity, i.e. xmod, the model-counterpart to the target quantity x. Typically, it is too difficult to directly calculate xmod because the equations are too complex to solve. Researchers hence resort to simulation and build a model that approximates the theoretical model so that the computer can retrieve numbers for the corresponding quantity xsim of the simulation model, see Figure 1 for a schematic display.

Scheme showing relations between the real world, modeling, simulation and experiments.
Up to this point, simulation helps to determine quantities, but these quantities are merely modeled versions xsim of the quantity xmod that itself has been modeled. Now the feedback loop displayed in Figure 1 comes into play. Scientists can vary the model input or parameters of the model and ‘observe’ how xsim changes. This is an experimental activity, but one that does not deal with nature or some material system in the laboratory, but rather with the simulation model.
Consider the elasticity example. The researchers are interested in properties like elasticity of the mixture of two substances. Of course, large data banks of substance properties are available, but they cannot possibly specify properties of all mixtures: since these properties change with mixing ratios, the possible entries are too numerous for data banks to ever be exhaustive. Typically, data about elasticity of the mixture of interest exist, but are sparse.
Moreover, no theoretical model will be able to match the measured data perfectly, because any particular mixture of particular substances will present a special case not fully covered by general theory. Adapting parameters at the right places of the model, however, allows matching the simulated values with the measured ones. If the theoretical model is good enough and if the parameters are adequately chosen, matching the available data will give confidence for predicting new data, designing processes, etc. Even if the theoretical knowledge has important gaps, and even if the software is imperfect, iterative adaptation does not care and might still provide a pragmatic solution. Thus flexibility, rather than sheer computational power, is the crucial property!
While models with adjustable parameters have been around much longer than computers, practical hurdles had limited their use in the past. The easy availability of computers and optimization software has tremendously lowered these hurdles. It has become much easier to utilize the adaptability of models, 4 and therefore much more tempting to succumb to the lure of making models fit by adjusting enough parameters. In terms of the tension described above: adjustment targets imitation rather than representation.
This general description of methodology suggests a certain organization of simulation. Adaptation creates a networked structure. The theoretical model is a shared object of a wider community. A generic simulation model adds discretization and parameterization, but leaves parameter values still unassigned. It is shared by a smaller group of researchers. Moreover, this group can use several of these models. This model still is a not-yet complete entity that is ready to be specified by a research team or working group. Facets of specification include particular implementations, choice of parameterization, tricks to speed up computation and formulating a mixing rule (in our example). Additionally, the measured data to which a model gets adapted exert a strong influence. Different teams of researchers might prefer to work with different sets of reference data. One reason is what data collaborating experimental groups can produce. Depending on reference data, specifications of simulation models will be different.
Simulation thus shows a peculiar embedding and dis-embedding pattern: the theoretical core of the simulation model remains invariant between different teams of scientists, but other specifications do not. Thermodynamic simulation models, for instance, differ greatly in performance depending on parameter assignments. The instrument undergoes a more significant transformation than suggested by ‘embedding’. It is apt, therefore, to remember that simulation has been characterized as an ‘instrument of instruments’, indicating an additional layer of freedom. This is a philosophical –and admittedly a bit abstract– argument showing the plausibility of my claim about how the methodology and organization of simulation are intertwined. Now, I return to empirical evidence.
Re-organizing the instrument
If adaptation and organization indeed play a major role, can one find traces of this fact in available data on simulation? This section gives some positive evidence. Up to the end of the Second World War, ‘computer’ was the name for workers, mostly female, who carried out elementary calculations (Grier, 2003). They were assembled in large groups for tackling voluminous calculation tasks by cumulating easy numerical steps in an organized way. The primary object to be computed were ballistic tables and the military had the resources to organize the computers. This ant-like business was radically reorganized when electronic computing machines became available. The digital computer has been borne as an elephant. As is well known, it had been developed during the Second World War, with one of the first uses in connection to the military-funded Manhattan project. The early machines, like ENIAC, were slow giants according to today’s measurements, but could easily compete with large colonies of human computers. The mainframe computer, hence, is elephant-like because it replaces so many human computers and because it required enormous funds and energy to develop this instrument.
It is interesting that simulation partly goes back to the same source, or at least simulation understood as digital computer simulation. Early simulation methods like Monte Carlo co-evolved with the computer. Other methods like finite differences had been around in numerical mathematics, but experienced a tremendous upswing when implemented on a computer. Simulation, therefore, is a sort of co-elephant to the mainframe computer. Both computer and simulation gain traction in science during the 1950s and 1960s. Both are paradigm examples of ‘big’ science, which significantly differs from ‘little’ science not only in numbers of researchers and quantities of funds but also in how science is socially organized. 5 In the useful terminology of Marcovich and Shinn (this issue), computer simulation belonged to a bureaucratic work environment.
And in fact, one can confirm this expectation by looking at data from the ISI Web of Science, a commercial databank that collects information about scientific papers and their citations. Figure 2 displays the course of simulation in a rough and ready way. For every fifth year from 1900 to 1985, it counts all published papers with the word “simulation” in title or abstract. This number by itself is not so significant, because the overall number of papers published per year increases with time. So the scale shows the percentage of papers on simulation among all papers of one year. I do not distinguish between papers with simulation as their subject and papers that use it as an instrument. Studying science by quantitative analyses has its shoals. One thing, however, is so obvious that I would trust the picture: up to around 1960, the percentage is consistently very low and from then on it shows quick growth.

Percentage of papers with ‘simulation’ in the title or abstract among all papers published in the ISI database 1900–85 (five-year intervals).
Although by 1965 the electronic computer had been available at least for one decade, there existed only a few big mainframe machines that attracted the attention of pioneers of simulation, but did not result in a discernibly high number of papers. Over the 1960s, simulation experienced its first turn, a turn that every reader might correctly think is straightforward, if not trivial. My main thesis about organization and simulation will be about a more unexpected second turn. First, have a look at the data. Figure 3 extends the data of Figure 2 to recent times, collected for every year from 1979 to 2010. There is another sharp bend at around 1990 that divides a period of slow growth from a period of much faster growth. This I call the second turning point of simulation.

Percentage of papers with ‘simulation’ in the title or abstract among all papers published in the ISI database 1979–2010 (one-year intervals).
Taken together, both figures show a first turn from standstill to slow growth around 1965, and a second turn from slow growth to fast growth around 1990. Both turns stand for a change in magnitude of about equal dimensions because the slope of the graph (speed of growth) changes by a factor of ten both times. The first turn is totally expected, but what happened in the second turn?
If computational power is the pivot factor for simulation, the second turn appears to be unlikely. Of course, computational power is an important element and arguably the key to the first turn of simulation. Right from the start, however, the power of computers increased quite steadily in time – Moore’s law expresses this fact. Why then should there occur a relatively sharp transition around 1990?
If one ties the conception of simulation to computational power, the 1990s turn comes as a total surprise. Some surprises teach how rich the world is in relation to our expectations. Others indicate our conceptions are misleading. The present subject, i.e. the 1990s turn, is of the latter kind, I argue. We should look with suspicion at the computational-power-conception of simulation.
There is a change that occurred around the 1990s and that potentially explains the second turn. It is the availability of lab-scale computers, i.e. computers that are networked and easily available like desktop computers, workstations or clusters of them in the basement of research institutes. They do not increase computational power compared to centrally maintained high-performance machines. On the contrary, they usually offer only second- or third-rank power. The point is they are part of a different regime of how simulation is organized.
My main claim is this: the 1990 turn signals a new quality of simulation in terms of its social and cognitive organization. Methodology and organization of simulation are interrelated and the broad success depends on their relationship. In the metaphorical language of the jungle, it is a turn from elephants to ants. In the metaphorical language of the history of science, it is a turn from big science to small science. In the terminology of Marcovich and Shinn (this issue), it is a change in the configuration of the instrument, namely a turn from bureaucratic to autonomous work environment.
Simulation, then, thrives on a particularly exploratory and iterative mode of modeling. This mode makes extensive use of the feedback loop in modeling highlighted above. Using this loop in an exploratory manner allows working with preliminary models that are adapted over the course of modeling, based on observed performance rather than theoretical specification. The exploratory mode presupposes low cost, both financially and in terms of the researcher’s energy, for changing, adapting, evaluating the models, etc. The key for such mode of modeling is that computers are easily available and networked – a condition satisfied from the 1990s onwards.
The exploratory and iterative mode of simulation does not go well with centrally controlled computers and expensive computing time, it is not yet clear how relevant this mode of modeling in fact is. Can it account for the impressive second turn? I am convinced it can. What could such an argument look like? I propose the following: I discuss a scientific field whose success is, first, building on the exploratory-iterative mode, second, coincides with the 1990s turn, and, third, this field is of sufficient import to simulation-based science that it can exemplify the turn rather than being some exception. This is what the next section will provide.
The argument – density functional theory
I want to argue about the organization of simulation using one particular example, namely quantum chemistry. Moreover, I will concentrate on one specific part of it: density functional theory (DFT). This case is typical for simulation, I will argue. DFT in fact satisfies all three conditions mentioned above. The way simulation became socially and cognitively organized in the 1990s contributed much to the rise of DFT. It is, according to my claim, a paradigmatic instance for the success of ant colonies.
Quantum chemistry: mirroring the first and second turns of simulation
Here is a fast-forward history of quantum chemistry. In the early 20th century, chemistry was firmly established as a discipline with a strong experimental culture, considered to be profoundly different from the rational–theoretical branch of physics. The difference was put into question when the new quantum theory was formulated in the 1920s. It describes the electronic structure of atoms and molecules, which in turn determines their chemical properties, like bond energies. Hence quantum theory seemed to establish a bridge between theoretical physics and chemistry. At least such a bridge started to look like a real possibility. The Schrödinger equation (1926) seems to provide a useful formulation of the problem, because this equation contains all information about the electronic structure of molecules. Shouldn’t one be able to compute chemical properties by solving the Schrödinger equation?
This question was answered in the positive very quickly. The 1927 joint paper by the German physicists Walter Heitler and Fritz London is widely acknowledged as the first seminal work in quantum chemistry. There they treated the simplest case, the hydrogen molecule, and argued that homopolar bonding could be understood as a quantum phenomenon. The argument was of mathematical nature: from the Schrödinger equation it follows that two electrons with antiparallel spin (a quantum concept) that aggregate between two hydrogen protons reduce the total energy, i.e. create a stable configuration.
It soon became clear, however, that moving beyond the simplest cases leads into terrain utterly intractable for reasons of computational complexity. Basically, the Schrödinger equation expresses the energy via a wavefunction Ψ(1,2,…,N) that has as variables all N electrons of an atom, molecule or bunch of molecules. The electrons interact and hence Ψ has 3N degrees of freedom (three dimensions of space, leaving spin aside), a number of discouraging cardinality in many circumstances. Practically, to solve Ψ is extremely difficult and computationally demanding – it can be seen as a paradigm of computational complexity. Quantum chemistry did not die out, but progressed in a ‘quasi-empirical’ mode where measured values replaced those quantities in theoretical models that were too complicated to compute. There exist comprehensive historical accounts of the history of quantum chemistry of which I would like to mention the chapter in Mary Joe Nye (1993) and, in particular, the recent book by Kostas Gavroglu and Ana Simões (2012).
These accounts portray how the pathway of quantum chemistry meandered between physics and chemistry and also between theoretical and experimental means. The historians agree that quantum chemistry has been established as a sub-discipline of chemistry due to the electronic computer. With computational and simulation methods one could tackle models that had been out of reach previously. In a sense, one elephant (quantum theory in the form of the Schrödinger equation) had got stuck, but when the second elephant (the computer) arrived, they together were successful. Numerical methods were developed that aimed at solving the Schrödinger equation ab initio, i.e. without recurring to measured values. The timeline coincides with the first turn to simulation. The historical accounts agree that the computer changed the game and superseded the quasi-empirical approaches (around 1970). After that the written histories of quantum chemistry fade out.
DFT: stunning success
There is more to discover, however, namely a turn that elevated quantum chemistry to one of the most visited spots in all of science. Though there is not yet a commonly accepted terminology, it is apt to call this turn the turn toward computational quantum chemistry. 6
All those readers who trust my line of argument will expect that this turn happened around 1990 and is related to simulation instrumentation and how it is organized. The remaining part of this section will convince you this expectation is correct.
One theory plays a particular role here, namely DFT. There are a couple of theories undergoing similar developments, but DFT arguably deserves the most attention. Among quantum chemists, to begin with, there is widespread agreement on the special role that DFT plays among a couple of ab initio methods: ‘The truly spectacular development in this new quantum chemical era is density functional theory (DFT)’ (Barden and Schaefer, 2000: 1415)
Let me first make evident that the development of DFT is in fact spectacular. After that, I will argue why and how this development is related to simulation as an instrument. As I mentioned, so-called ab initio methods flourished in quantum chemistry from the 1970s onward. All these methods made heavy use of the computer and numerical approximations. Two prominent examples of ab initio methods are Hartree-Fock and coupled cluster. DFT has its origins, in the 1960s, in condensed matter physics and was an influential theory in physics since then, but marginalized in chemistry. The ISI Web of Science database consistently counts only around 30 papers per year that were published on DFT 7 in all of chemistry (according to the disciplinary assignment of ISI) throughout the 1970s and up to the late 1980s. However, around 1990 a tide change happened. Figure 4 displays this development vividly and makes it evident that DFT underwent a serious phase transition in the 1990s. Between 1990 and 2005, the number of scientific papers on DFT grew exponentially from around 30 to a level of more than 4,000 per year. Moreover, the fast growth has continued since then. Today, more than 15,000 papers on DFT appear each year in the sciences.

Number of papers with ‘density functional theory’ in title or abstract per year in the ISI Web of Science database.
I would even dare to claim the staggering growth of DFT has made it the dominant scientific theory of present times. In addition to the number of papers published, I want to add another piece of quantitative evidence. Consider the journal Physical Review, a flagship journal in physics and chemistry. What are the most highly cited papers that appeared ever in this journal? As you might expect, the paper by Einstein et al. (1935) is high on this list. The same is true for the one by Bardeen et al. (1957) on superconductivity, or the one by Binning et al. (1986) on the scanning tunnel microscope. These are well-known ‘elephants’ with high impact not only in science, i.e. physics in this case, but also in history and philosophy of science.
The breaking news is that among the top ten most highly cited papers more than two thirds are about DFT. 8 Yet, DFT does not play a role in either history or philosophy of science. Given the kind of evidence presented, the missing attention from the side of science studies seems bewildering. In my opinion, this is an instance where the ants escaped notice though they actually outweigh the elephants. Counting citations is a bit like weighing ants and the question is how significant such quantities are. A critical reader might ask whether and how the turn of DFT and its recent success is related to simulation instrumentation in a systematic way.
A closer look on DFT
Answering this question requires a closer look at DFT. What is DFT about? Quantum chemistry deals with the electronic structure of atoms and molecules. The Schrödinger equation expresses this structure as a wave equation of frustratingly high complexity. Much of quantum chemistry was and still is occupied with finding suitable numerical approximations of what this equation implies about certain values of chemical interest.
In a nutshell, DFT is a theory of the very same electronic structure that circumvents the problem of solving the Schrödinger equation. DFT expresses the energy in a different way, namely in terms of the (joint) electron density – roughly the more likely it is that electrons visit a certain location in space, the higher the electron density. The density hence is an object in space and has only three degrees of freedom. On the face of it, this formulation reduces complexity starkly.
The computational advantages of such a reduction of complexity brought this approach into heuristic use in engineering fields, even before the theory was formulated. The theoretical condensed matter physicist Walter Kohn played a major part in advancing this approach to the level of theory. He and his colleague Pierre Hohenberg contributed the two founding theorems (Hohenberg and Kohn, 1964). They established that the corresponding electron density indeed uniquely determines the ground state energy, that is, the energy is a function only of the electron density and thus can be calculated without reference to the Schrödinger equation, at least in principle. However, the promise is one in principle, whereas in practice one gets from the frying pan into the fire.
Note that the reported 1964 results did prove that there exists a function f that gives the energy and that is dependent only on the electron density. 9 While the energy entirely depends on the exact form of this function, the theorem does not give any clue as to how that function looks or how it can be determined. The space of mathematical functions is extremely large, definitely larger than a haystack, hence, to actually determine one particular function might be very difficult. As long as the function is unknown the advantage of reduced complexity does not pay out. Kohn was aware of this shortcoming and in the following year he introduced, together with his co-worker Liu Sham, a practical computational scheme (Kohn and Sham 1965) for approximating the desired functional. This scheme (counterfactually) postulates an idealized situation that makes the approximation of the unknown functional computationally feasible.
The mentioned 1964 and 1965 publications were – and still are – very influential papers. One can read this from bibliometrical evidence. Indeed, they lead the list of most highly cited papers in Physical Review. Eventually, in 1998, Kohn received the Nobel prize ‘for his development of density functional theory’. The reader might wonder how this story about theoretical physics in the 1960s can possibly throw light upon the 1990s turning point of simulation.
Well, I want to ask for only a little more patience. The next step is the observation that the Kohn–Sham scheme for density functionals is attractive because it is computationally relatively ‘cheap’, but it does not provide results of high accuracy. This is sufficient when the objects of interest are regular structures like in crystallography, but most of the interesting questions in chemistry require higher accuracy. DFT therefore was of little use in chemistry, which resonates with the observation made previously that only few papers in chemistry dealt with DFT. The situation is reflected in the following statement by chemists:
However, the correct functional of the energy is unknown and has to be constructed by heuristic approximation. Initial functionals, based principally on behavior of the electron gas … [the Kohn–Sham scheme], were lacking in the accuracy required for chemical applications. Breakthroughs over the past two decades … have led to the development of functionals capable of remarkable accuracy and breadth of applicability …. (Friesner and Berne, 2005, 6649)
10
Simulation re-organization
What brought the breakthrough in accuracy around 1990 that opened the doors to chemistry for DFT? The answer makes use of the results obtained in the section [Computer simulation as instrument’ above, where the feedback loop in modeling was highlighted as a crucial element in simulation modeling. Some feedback loop always plays a role in modeling, since hardly any model will be entirely correct from the beginning and hence needs to be modified over the process of model building. It is a different thing, however, if modelers deliberately leave their model unspecified and open to adaptation.
In the case of quantum chemistry, ab initio modeling was targeted at numerically solving the Schrödinger equation, The success of DFT came with a new strategy that hybridized ab initio and semi-empirical modeling. Researchers started with a plausible functional (the theoretical, ab initio part) that included multiple parameters open for assignment so that the behavior of the functional could be made to match known data (the semi-empirical part).
The new accomplishments around 1990 have been multi-parameter functionals. They do not stand for a breakthrough in theory, but rather for a breakthrough in the potential of adaptation. Researchers could use such functionals for following a strategy that thrives on the feedback loop of modeling. The strategy was new in that the starting point was not any longer the best available model, but a highly adaptable one. This strategy could become an integral part of the simulation instrument only around 1990. Exploratory modeling of this kind requires iterating modifications of the model (especially its parameters) for sounding out model behavior. This in turn requires cheap and easy access to computing technology – a condition that was in place when desktop computers, workstations and other microcomputers quickly spread around 1990.
If this point is valid, it would suggest that DFT does not thrive by developing functionals ever closer to the correct one. Instead, DFT-based approaches should splinter and develop a variety of functionals, each one being capable of adapting to different types of substances and chemical and physical conditions. This is indeed what happens: there is a flurry of functionals implemented in more than hundred software packages. Manuals even advise researchers not to trust single functionals, to try out several and to combine them into so-called hybrid functionals. Since there is no known reason why one particular functional should work best, researchers have to explore by experimentally fitting various functionals to their data. Typically, working groups and laboratories have a laundry list of which functionals are promising candidates for adapting them to certain types of substances and materials.
Thus, the astounding upswing of DFT since around 1990 is happening on the basis of a theory already developed in the mid-1960s. DFT thrives on the plurality of functionals and a large variety of adaptation practices. Together, they combine accuracy with feasibility in a way that makes DFT so attractive. In other words, the 1990s turn is not based on a general theoretical breakthrough, rather on very many particular achievements. At this point, simulation methodology and organization are intertwined. In a sense, DFT-researchers are organized like a colony.
One institution that is especially ill equipped to account for ant colonies is the Nobel committee. Of course, the overwhelming success of DFT in chemistry and related fields could not be overlooked. Hence Walter Kohn received the Nobel prize in 1998 for DFT, although he was a theoretical physicist and got the price – to his own surprise – for chemistry. Kohn shared the prize with John Pople, a mathematically minded chemist who had a leading role in promoting computational modeling in quantum chemistry. Pople was rewarded ‘for his development of computational methods in quantum chemistry’. 11 Pople was one of the pioneers in the 1970s who realized early that software would be key for fostering quantum chemistry. He and his colleagues did much for standardizing procedures and compiling the software package ‘Gaussian’, which is still the market leader today. However, Gaussian did not include DFT methods (up to the 1990s) because of their lack of accuracy in chemical problems.
What makes DFT researchers look like an organized colony? Ant colonies quickly spread out and explore new territory on very many pathways. At the same time, sustaining this undertaking requires keeping the colony connected. DFT shows a similar pattern of organization. A great number of small teams of researchers locally adapt a number of simulation models to their particular environment. These teams typically operate in a sort of niche, i.e. are interested in particular properties of particular materials under particular conditions. In the easiest case, such groups are successful adjusting and re-combining existing simulation models. On today’s networked computers, a great number of models is easily available and it is a straightforward task to take a number of them and test which one fits best to the situation of interest. Notably, this kind of activity calls more for computational scientists of a general stripe than experts in quantum theory. Arguably, the number of users of DFT is much higher than the number of quantum theoreticians.
Another case is when a small team or individual person develops a new parameterization, for instance a new multi-parameter functional. This functional then easily can be made available to other local teams that might want to include this functional among their candidate models, even if they only want to use it for showing the relative superiority of their own favorite model. Another factor is the networked infrastructure on which models and data can quickly travel; more than hundred software packages are available that contain DFT methods. Different groups develop very specialized types of functionals for particular materials. On the existing networked infrastructure functionals from different working groups can quickly be exchanged via software plug-ins and even re-combined as hybrids, which then create new functionals. One can see this process as a variant of dis-embedding and re-embedding, in the terminology of Joerges and Shinn (2001).
Simulation is an instrument that crosses the boundaries of local research teams as easily as that of scientific disciplines. Again DFT is a telling example. For decades, it was firmly rooted in condensed matter physics, but since the 1990s turning point chemistry has the largest share, weighed by the disciplinary assignment that the ISI database gives to articles written on DFT. Though chemistry has the largest share, there is a multitude of other disciplines, from engineering to materials science, that also have significant shares. If a scientific discipline is indicated by a color, DFT has become a multicolored object.
Overall, DFT is an important instance of simulation in recent science. It shows a remarkable success that starts in the 1990s, exemplifying a general and significant turn for simulation. The somewhat lengthy chain of argument has established that explaining the 1990s turn and the following success of simulation depends on how simulation is organized as an instrument.
Conclusion
I want to close by stepping back from the case of DFT and widen the perspective to simulation in general. Here is a suggestion and a challenge. The suggestion is about what Marcovich and Shinn (this issue) call instrument configuration. They introduce and discuss three different configurations and I would like to suggest that the 1990s turn for simulation introduced a fourth configuration. Before, simulation had been more on the bureaucratic side, organized by centrally maintained computing centers. This type still exists in today’s supercomputing centers that employ a large number of people and conduct rigorous planning of how their expensive computing time is used. After the turn, however, the instrument is expansionist on the basis of re-combining and adapting autonomously and locally developed components. The organization of simulation achieves something that can be called instrument multiplication. Simulation is a very special instrument because of the weak conditions materiality imposes. A code of 100 lines can travel as swiftly and can be adapted (nearly) as easily as a code of 1 million lines. Hence simulation escapes the tension Marcovich and Shinn observe between massification and diversity – it exemplifies both. The configuration is autonomous, expansive and diversifying.
The challenge is the problem of sprawl and consistency. It is the downside of simulation’s success. If this success is based on adaptation work at the dispersed fringe of a discipline or theory, like developing and adapting density functionals, then the internal consistency of the entire undertaking is questioned. Granted, simulation is configured as suggested above, i.e. autonomous, expansive and diversifying. How can internal consistency be maintained? In the case of DFT, the problem is to ensure that functionals designed for different niches agree in their predictions of test cases. The simulation community itself has recently discussed this concern with the ‘reproducibility of density functional theory calculations’ (Lejaeghere et al., 2016). The proposed solution is exploring and testing a number of model families and confirming statistically they give consistent results for reference data. In other words, the problem is addressed by seeking an imitation approach to a reference question – a truly simulationist answer.
Footnotes
Acknowledgements
This article profited from the critical audiences at the workshop on scientific instrumentation (Paris 2016) and at the Bielefeld I2SoS colloquium. I am particularly thankful for helpful suggestions by Ann Johnson, Anne Marcovich and Terry Shinn.
Funding
Work on this article was supported financially by DFG priority program SPP 1689.
