Abstract
Probability theory is at the base of modern concepts of risk assessment in mental health. The aim of the current paper is to review the key developments in the early history of probability theory in order to enrich our understanding of current risk assessment practices.
Introduction
What events should we fear? What weight should we attach to our fears and what measures to protect others and ourselves are proportionate? These are age-old and universal questions. Prior to the sixteenth century, ‘fate’ would have played a large part in any consideration. Today, an empirically based answer in any field of human endeavour can only be framed in terms of probability theory and risk assessment. Increasingly, clinicians in mental health are also expected to use risk assessment in their approach to reducing the harms experienced and committed by their patients. These assessments include estimations of the likelihood of violence and self-harm. However, current methods of risk assessment are very diverse and include clinical risk assessment based on the opinion of an experience clinician, consideration of a list of factors that are thought to pose a risk to the patient, actuarial-type risk categorization based on a score derived from a risk assessment scale, or a combination of these methods, sometimes referred to as structured professional judgement (Undrill, 2007).
This narrative review outlines some important historical aspects of probability theory that have a bearing on contemporary approaches to risk assessment in mental health care. Acknowledging that the term risk assessment has diverse meanings in mental health, the paper focuses on assessments of the likelihood of discrete future events such as acts of violence or suicide. The first part of the paper is an outline of the history of probability theory from its origins in the sixteenth century to the emergence of Chaos theory in the twentieth century. The second part is a discussion of how the ideas of Randomness, Loss, Accuracy, Certainty, Belief, and Chaos, developed by the pioneers of probability theory, remain relevant today for clinicians undertaking risk assessments.
A brief outline of probability theory
Surprisingly, a mathematical approach to the possibility of future harm eluded Greek and Roman thinkers, and probability theory emerged in Europe between the sixteenth and twentieth centuries. Today, a unified body of statistical knowledge is applied on a daily basis to such diverse fields as astronomy, aviation, climate science, economics, insurance, medicine, meteorology, political science, psychology and superannuation.
Blaise Pascal (1623–62), often cited as the original pioneer of probability theory, is perhaps best remembered for ‘Pascal’s wager’ – that a person should live on the assumption that God exists, even though this cannot be proved, because the value of eternal life is so great and the cost of damnation so severe that a rational person should simply wager that God exists. Pascal is also linked to more secular expressions of risk, for example he is known to have sponsored the monks of the Port-Royal Monastery in seventeenth-century France, where a ‘best seller’ of the era, La Logique, ou l’art de penser originated (Arnauld and Nicole, 1662/1964). In the final chapter of La Logique, entitled ‘Belief in future contingent events’, the monks used the example of lightning strikes and wrote that ‘Fear of an evil ought to be proportionate not only to the magnitude of the evil but also to the probability of occurrence’ (p. 361). This combination of a probability and loss is one of the earliest and most clear definitions of risk.
Despite the explicitly numerical nature of their definition of risk, the monks provided no guidance as to what exactly probability was or how it could be determined. In recent years, primary source documents from the very early history of probability theory have become readily available. These include: the correspondence between Pascal and Pierre de Fermat (1601–65) (Pascal and Fermat, 1654), and between Jacob Bernoulli (1654–1705) and Gottfried Wilhelm Leibniz (1646–1716) (Bernouilli and Leibniz, 1703); Abraham de Moivre’s (1667–1754) treatise on gambling, De Mensura Sortis (On the Measurement of Chance) (de Moivre, 1711/1984); and the monumentally important, yet obscure, text by Thomas Bayes (1701–61), with its explanatory forward by Richard Price (1723–91) (Bayes and Price, 1763). Furthermore, the rich and fascinating history of the emergence of probability theory and the lives of its key contributors have been brought to life by histories written from the perspectives of philosophy (Hacking, 1975) and mathematical scholarship (Gullberg, 1997; Stigler, 1986, 1999), and from the wish to popularize important scientific ideas (Bernstein, 1996; McGrayne, 2011; Mlodinow, 2009; Taleb, 2004).
Randomness and equity: Cardano, Pascale, Fermat and de Moivre
Gerolamo Cardano (1501–76), Pascal, Fermat and de Moivre pioneered the mathematics of games of chance and can be considered the originators of probability theory (Hacking, 1975; Gullberg, 1997). Although there is little doubt that each hoped to apply probability theory to more complex problems than gambling (Stigler, 1986), between them they explored and solved probability problems where the chance of the outcomes can be known with certainty. They considered games with coins, dice, cards and, had they lived in current times, might have considered roulette wheels, electronic lotteries and gaming machines.
By all accounts, in 1525 Cardano wrote the first numeric expression of the chance of a variety of outcomes (see Cardano, 1663). He was a physician and professional gambler, known for having conceived of negative numbers (Bellhouse, 2004; Hacking, 1975). Although a prolific writer on diverse subjects, Cardano’s monograph on gaming, Liber de ludo aleae, was not published until 1663, many years after his death. This may have been partly due to a wish to keep his edge over those with whom he gambled, as suggested by the presence of a chapter on cheating. As a result of this delay, the origin of probability theory is usually attributed to Pascal and Fermat, who corresponded on the topic between 1654 and 1660. Pascal, a deeply religious mathematician, initially wrote to Fermat, a lawyer and respected mathematician 20 years his senior (now most recognized for his tantalizing and until recently unsolved ‘last theorem’). Pascal suggested a solution to the apparently simple problem of how to divide equitably the stakes if a game of dice is interrupted (Pascal and Fermat, 1654). Their correspondence, while somewhat incomplete, ends to their mutual satisfaction after a total of 10 brief letters. In these letters, they agree that probability can be expressed numerically, concur that the price of the game or risk is a fraction of a resulting loss, and calculate probability by considering all equally likely outcomes. In order to do this they agreed on the need to multiply the probabilities where one outcome is dependent on the other, such as the chance of rolling two sixes in a row (1 in 36) and on how to calculate the odds of rolling one six in two throws (11 in 36). This simple second calculation, and the numerous extensions of this type of problem, had eluded earlier thinkers. In their letters they realized that the 6 in 36 chance that one die will show a 6, and the 5 in 36 chance that the other will turn up a six when the first was a 1, 2, 3, 4 or 5 can be expressed as simple analysis of branched outcomes. This analysis of branched outcomes provides a model for calculating probabilities where the possible outcomes of each trial are equally likely. It can also be applied to more complex cases where not all the outcomes are equally likely, but where each can be assigned a certain value. It is clear from their letters that Fermat and Pascal understood the wider importance of their findings and the possibility of generalizing their analysis beyond games of chance. However, neither wrote much more on probability, although each made numerous other contributions to mathematics and the natural sciences.
In his monograph De Mensura Sortis (The Measurement of Chance), de Moivre (1711/1984)extended the work of Fermat and Pascal to other games of chance, modifying the notation and dealing with increasingly complex and more general cases (Gullberg, 1997). It is in de Moivre’s work that the familiar and modern definition of risk emerges in a numeric form: R = L multiplied by P, where R is the risk, L is the loss, and P is the probability of the loss (de Moivre, 1711).
Despite his achievements, de Moivre himself was dissatisfied with one important aspect of his work. He knew that, in most practical problems, the set of outcomes was not equally likely and could not be easily estimated (Stigler, 1986). He believed that with a sufficiently repeated number of trials of any event, the proportion of events to non-events approached the probability of its occurring in a single trial. However, he had no proof of this and no method for calculating just how many trials were needed to have any defined degree of certainty in the outcome of future trials. For example, if you had no knowledge of a physical die, but were given a string of numbers generated in this way, you would have no way of knowing how many numbers you would need to observe in order to be sure that 1, 2, 3, 4, 5 and 6 were equally probable and that the average of the numbers in the string was 3.5. This problem became known as the problem of ‘inverse probability’ and is now also referred to as the probability distribution of an unobserved variable. The problem of a risk assessment in mental health is a typical example of an inverse probability problem because, although the violent outcomes may be all too evident, we cannot directly observe the causes of violent behaviour. Inverse probability problems are similar to most problems faced in the real world and differ greatly from the problem of gambling where the machines of chance such as dice and roulette wheels can be observed directly. Further, the interpretation of inverse probability lies at the centre of an unresolved question in mental health risk assessment: to what extent is it valid to generalize the findings from group data gathered by repeated observation to inferences about individual cases.
Accuracy: Bernoulli and Leibniz
Jacob Bernoulli was one of five members of the Bernoulli family to have contributed to probability theory. Prior to Bernoulli, statistics was concerned with the results of games of equal or, if not equal, known chance. It could, for example, settle the probability of drawing a particular combination of red and green marbles from a jar with a known number of each type. Bernoulli took serious interest in the inverse probability question: how, by simply observing the colours of repeatedly drawn marbles, you could determine the ratio of colours of the marbles in the jar. He concluded in a letter (3 Oct. 1703) that ‘what is hidden from us by chance a priori, can at least be known by us a posteriori, from an occurrence observed many times in similar cases’ (Bernouilli and Leibniz, 1703). Bernoulli was also the first to express probability using the modern nomenclature of a decimal between 0 and 1. These developments appeared in his best-known, yet posthumous, book Ars Conjectandi (The Art of Conjecturing) of 1713. Here, Bernoulli demonstrated his proof of what is now known as ‘the Law of Large Numbers’ or the ‘Central Limit Theory’. In its simplest form, the Law of Large Numbers suggests that the results of a large number of trials will be similar to the probability in a single trial. To some, the Law of Large Numbers seems little more than common sense, but Bernoulli proved mathematically that the results obtained in a larger number of trials is closer to the expected value than the results obtained after a smaller number of trials. On completing his proof, Bernoulli wrote to Leibniz on 3 Oct. 1703: even the stupidest man knows, by some instinct of nature per se and by no previous instruction, that the more observations there are, the less danger there is in straying from the mark, it requires not at all ordinary research to demonstrate this fact accurately and geometrically. (Bernouilli and Leibniz, 1703)
Although Leibniz accepted Bernoulli’s proof, he was less enthusiastic about the implications of the Law of Large Numbers. In his reply, Leibniz can be thought of as starting an enduring controversy about the limits of applying probability theory to natural occurrences when he wrote the following (3 Dec. 1703).
When we estimate empirically, by means of experiments, the probabilities of successes, you ask whether a perfect estimation can be finally obtained in this manner. You write that you have found this to be so. There appears to me to be a difficulty in this conclusion: that happenings which depend upon an infinite number of cases cannot be determined by a finite number of experiments; indeed, nature has her own habits, born from the return of causes, but only ‘in general’. (Bernouilli and Leibniz, 1703)
While Bernoulli appears not to have been greatly deterred by the comments of the older and more accomplished Leibniz, he was disappointed by the sheer number of trials required to reach a ‘perfect estimation’ and was aware of his failure to elucidate a method for knowing the degree of certainty one could have according to the numbers of observations. De Moivre had provided an approximate answer to this question, but it was left to Bayes, Price and Simon Laplace (1749–1827) to settle this measure of certainty.
Certainty: Bayes, Price and Laplace
Bayes was a Presbyterian minister in rural England. He published little during his life, but understood the new field of calculus and was probably familiar with the work of de Moivre (McGrayne, 2011). Historians are unaware of the exact date, but some time in the 1750s Bayes considered the following problem: ‘Given the number of times in which an unknown event has happened and failed: Required the chance that the probability of its happening in a single trial lies somewhere between any two degrees of probability that can be named’ (Bayes and Price, 1763: 376, original italics). By this he meant that, given simply observed events and non-events, how certain can we be of our judgements about future events? To this important problem Bayes provided an elegant but esoteric proof by imagining the position of balls on a table when all that was known was whether the balls fell to the right or left of an imaginary line. He then filed away his proof with other writings, apparently telling no one. On his death, Bayes’ papers were left to the executor of his estate and fellow mathematician, Price. Price presented Bayes’ theorem, with an explanatory foreword, to the Royal Society in 1763. Price was no mere messenger; it has been suggested that Bayes and Price should be jointly acknowledged for the origin of Bayes’ theorem (Stigler, 1986). Price was also a preacher, widely known at the time in Britain and the USA as a social and political radical. He was friendly with Thomas Jefferson and John Adams and was a vocal opponent of the British war in America. He was also a formidable and practical mathematician, who applied what was known about probability theory to the fledgling insurance industry (Hacking, 1975). However, the importance of the work of Price and Bayes lay unexplored for half a century, until Laplace generalized Bayes’ findings and developed Bayes’ theorem in a form that is recognizable today in Essai philosophique sur les probabilités (Laplace, 1814). It has been argued that eponymous mathematical theorems such as Bayes’ theorem are invariably falsely attributed (Strigler 1999). In the case of Bayes’ theorem, it was Laplace who developed it in its most recognizable form: (P (A/B) = [P (B/A) P (A)]/P (B)), where P(A/B) is the posterior probability given the contingency of B, P(A) is the prior or initial degree of belief (the assumed base rate) and P (B/A) / P(B) is the relationship between B and A . Laplace also developed the concepts of sensitivity (the proportion of true positives detected) and specificity (the proportion of true negatives) and the concepts of the positive predictive value (PPV) (the proportion of true positives among test positive cases) and negative predictive value (NPV) (the proportion of true negatives among test negative cases). It was from the concepts of sensitivity and specificity that the receiver operator curve (ROC) was derived during World War II (Green and Swets, 1966). The ROC is a plot of the trade off between sensitivity and specificity. The area under this curve (AUC) has become a key measure of the ability of risk assessment instruments to discriminate between high and low risk groups (Mossman, 1994). Hence the now widely used measures of an actuarial risk assessment; the sensitivity, specificity, AUC and the PPV have a long and well documented history.
In Bayes’ terms, the PPV and NPV are direct expressions of the degree of confidence one can have in a contingent probability (such as when a risk factor is present) given an anterior probability – an assumed incidence or base rate. In fact, the most common expression of Bayes’ theorem (P (A/B) = [P ( B/A) P (A)]/P (B)) is an expression of PPV.
These statistics are now central to medical diagnostics. For example, they inform us that the likelihood of the carriage of HIV in a low-risk person with a positive antibody test is lower than in a high-risk patient with a similar test result, because the base rate (in Bayes’ terms, the P(A) or prior belief) is lower. When we undertake a violence risk assessment, Bayes’ theorem is at the centre of our considerations. Any violence risk assessment instrument (with a given threshold of sensitivity and specificity) will define a high-risk category with a higher proportion of those who will actually go on to be violent when applied to populations with a higher base rate of violence than when applied to populations with a lower base rate of violence. Hence, the interpretation of any risk assessment must involve an estimate of the base rate. This base rate, which is never known with complete certainty at the time of the assessment, is a Bayesian ‘prior probability’. This sort of estimate, present to a greater or lesser degree when any diagnostic test is applied in medicine, is in effect an educated, if highly educated guess. In a violence risk assessment, our degree of belief in the accuracy of the assessment ought to be tempered by our prior knowledge of the expected base rate. Hence, in Bayesian terms, the risk assessment serves not to give us an answer as to the probability of a future harm, but serves to modify our prior belief in the propensity for violence in the person being assessed.
It is then ironic that, although Laplace clarified how we could measure the certainty in processes with a random element, he was a strict determinist. In his book on probabilities, he imagined all of nature to be like a clock or the then current notion of a planetary system when he wrote: We ought then to regard the present state of the universe as the effect of its anterior state and as the cause of the one to follow. Given for one instant an intelligence which could comprehend all the forces by which nature is animated and the respective situation of those beings who compose it – an intelligence sufficiently vast to submit these data to analysis – it would embrace in the same formulae the movements of the greatest bodies of the universe and those of the lightest atom, for such an intellect nothing would be uncertain and the future as the past would be present to its eyes. (Laplace, 1814/1951: 4)
This passage has since come to be known as that describing ‘Laplace’s demon’. In expressing this, Laplace foreshadowed the more recent and more famous statement, ‘God does not play dice’, widely attributed to Einstein’s reaction to Heisenberg’s account of statistical mechanics. Laplace and Einstein both believed that all uncertainty about future events can be attributed to a lack of knowledge (epistemological uncertainty), and none to random processes (aleatory uncertainty) – a topic to which I shall return shortly.
Frequentist interpretations: Fisher
As a consequence of the work of Pascal, Femat, Bernoulli, de Moivre, Bayes, Laplace and others, by 1800 it was possible to describe the mathematics of well-known games of chance, to be secure in estimates of probability through repeated trials where the generators of chance are unknown or unobservable and, given certain assumptions, to calculate the certainty of inferences about the likelihood of future events after observing no more than their previous occurrence and non-occurrence. After 1800, Karl Friedrich Gauss (1777–1855), Francis Galton (1822–1911), Karl Pearson (1857–1936), Ronald Fisher (1890–1962) and a host of others advanced statistical knowledge in ways that cannot be easily summarized. Here I wish to highlight perhaps the most significant controversy in the history of probability theory: whether a probability can only ever be considered to be a proportion of observed events after a sufficiently large number of trials or whether a probability can be considered to be the degree of truth of a statement – as implied by Bayes and Laplace and as stated explicitly by John Maynard Keynes in his book A Treatise on Probability (Keynes, 1923). By the mid-1930s, the world’s pre-eminent statistical expert was Ronald Fisher. He had published the highly influential Statistical Methods for Research Workers (Fisher, 1925), which heralded the application of statistical methods originally applied by him in agricultural settings to other fields, including psychology and medicine. Fisher introduced the term ‘probability’ in relation to a p-value, meaning probability of obtaining a test statistic at least as extreme as the one that was actually observed, for example, by using his Exact test to determine if the likelihood that a difference in two proportions was not due to chance. Fisher also held strongly that probability was no more or less than a chance, such as an observed proportion in a large number of trials. He rejected any notion that a probability was an inherent quality or tendency for the observed events and believed that probability should not be regarded as an expression of the degree of truth of a statement.
Fisher’s strictly numeric view of probability as no more or less than the presently observed fraction derived from events and non-events has been called a ‘Frequentist’ view of probability. It can be contrasted with that of Bayes and Laplace who believed that probability, though expressed mathematically, was an expression of the plausibility of a particular belief. A Frequentist view of a risk for violence would, if sufficient data were available for the class of people assessed, be expressed as a simple fraction. By contrast, a Bayesian view recognizes the limitations of the available information, and uses the assessor’s prior knowledge (for example, of the base rate of violence) to influence the degree of confidence in an estimation of risk. As such, a Bayesian is less interested in determining a particular p-value for observed events, but seeks to measure, given available knowledge, the plausibility of an event. In the last 100 years, the importance and credit given to Bayesian ideas versus Frequentist conceptualizations has waxed and waned (McGrayne, 2011). However, in the social sciences and in mental health care, the degree of uncertainty in belief is often high. Few believe, as is required by a strictly Frequentist interpretation of probability, that a population mean is real, but simply unknown. In the social sciences, Bayesian interpretations are well accepted and it is assumed that the best one can do is to construct a credible confidence interval centred around the mean, after considering all the data and the assessors’ prior beliefs. The difference between these interpretations of probability might be significant in considering a violence risk assessment. A risk assessment with the purpose of modifying one’s prior beliefs, in other words, a risk assessment known to contain a subjective element, as suggested by Bayes, can be contrasted with a risk assessment that derives a numerical expression of the probability of future harm. These two views of probability as applied to mental health risk assessment might be interpreted quite differently as discussed below.
Chaos: Poincaré, Lorenz
In the last 50 years dispute about the nature of probability, whether Bayesian or Frequentist, appears to have lessened (McGrayne, 2011). Arguably this is because, whether it be an estimate of a true mean or a degree of confidence of a belief, in many cases the sample mean is not so informative about future events, as Leibniz declared in a letter of 3 Dec. 1703, ‘indeed, nature has her own habits, born from the return of causes, but only in general’ (Bernouilli and Leibniz, 1703, original emphasis). In the twentieth century, few would share Laplace’s view that an all-knowing all-thinking ‘demon’ could predict the course of future events, and most would believe that chance plays a major role in uncertainty. Belief in a deterministic universe has been undermined by the findings of quantum mechanics and particle physics, suggesting that, at a subatomic level, any method of observation affects what is being observed. What is perhaps better understood, but is less well known, is that profound uncertainly is a characteristic of some relatively simple mechanical systems, particularly those with a repeated feedback loop and interacting events. Chaos theory shows that when these conditions are present, for example in non-linear systems such as weather, economies and, in all likelihood, the multiple interacting neural networks in the human brain, the past cannot always be an accurate guide to the future, and gathering increasingly detailed knowledge of initial conditions will not necessarily result in more accurate predictions.
The first hint of what is now known as Chaos Theory came from astronomy. Difficulty with calculating the orbits of three or more mutually interacting bodies in orbit, for example the sun, the earth and the moon, was a well-known problem of classical mechanics in the eighteenth century. Its solution was the subject of a valuable prize offered by Oscar II of Sweden. Jules Henri Poincaré (1854–1912) was the first recipient of the prize – not for solving it, but for showing that it could not be solved without approximations. In fact, Poincaré demonstrated that many cases of the three-body problem do not result in stable orbits (Barrow-Green, 1997).
A better known example comes from Edward Lorenz (1917–2008), an American meteorologist who coined the term ‘the butterfly effect’. While the title of his 1972 paper, ‘Predictability: does the flap of a butterfly’s wings in Brazil set off a tornado in Texas?’ might be overblown (Lorenz, 1972), he did find that an accidental rounding error from 0.506127 to 0.506 (a 0.025% change) in initial settings of an early computer weather simulating program brought about totally unexpected and dramatic changes in the long-term forecasts (Lorenz, 2001). To this extent, the question of whether uncertainty in future events is a result of chance, or is born of a lack of knowledge of initial conditions, and the extent to which we perceive chance and risk as a numerical fraction or as a judgement about future events, become less relevant because (a) increased knowledge of initial conditions might not make very much difference to the accuracy of predictions, and (b) in practical terms, in complex systems the difference between uncertainty due to a lack of knowledge (epistemological uncertainty) and uncertainty resulting from truly chance events (aleatory uncertainty) cannot easily be distinguished (Taleb, 2007).
Six considerations for risk assessors from the early history of probability theory: randomness and equity, loss, accuracy, certainty, belief, and chaos
Risk assessment, randomness and equity
It is sometimes forgotten that a violence risk assessment is based on a probabilistic notion and is not really about the prediction of particular violent acts by individual patients. It is, rather, about estimating the incidence of violence in low- and high-risk groups. Even if one were to believe that that no process is truly random, it ought to be conceded that the number of intervening variables is so large that, in practical terms, individual outcomes are best regarded as a matter of probability. Nonetheless, evidence abounds that authors often slip into the language of prediction when discussing risk assessment. A search of articles indexed in Medline in the last five years reveals numerous papers with titles that include the word ‘violence’ and the term ‘predictive validity’. Almost invariably, these papers use the area under the curve (AUC) as a measure of ‘prediction’ or ‘predictive’ or ‘predicting’. This is surprising, because the authors of these papers almost certainly know that that AUC is not a measure of prediction, but is a measure of effect size of the association between risk factors and the violent events. Even when the AUC is high and there is a strong association between a risk score and violence, this does not mean that those who are categorized as at high risk are highly likely to commit a rare act of more serious violence (Large, Ryan and Nielssen, 2010; Szmukler, Everitt and Leese, 2011). Further, while the strength of an association or effect size is at the centre of risk assessment and probability theory, ‘prediction’ of the occurrence, type, place and timing of violence is quite literally the the stuff of science fiction (Dick, 2002). In reality, no violence risk assessment can determine what a person will do or where or when the violence will occur. This is well understood in other areas of risk assessment – insurance companies make no attempt to predict which individuals will have their car stolen, when a flood will happen or which doctor will face a malpractice suit, but instead they seek to group individuals with similar risk profiles into groups that are charged a similar premium. Moreover, the business model of insurance companies relies on a lack of prediction, for if a person knew with certainty that their belongings were safe, then why insure them? If an insurance company knew a loss was certain, the premium would have to exceed the replacement value. What insurance companies do try to estimate is the proportion of events among those insured in different risk categories. This is in essence the positive predictive value of PPV, which can be regarded as the true measure of the accuracy of high-risk categorization when measured against the eventual harms (Large, 2010).
What the positive predictive value shows is that, for any rare and serious harm such as a homicide or a suicide, the predictive value of risk assessment is invariably very low (Large, 2010; Large, Ryan, Singh, Paton and Nielssen, 2011; Szmukler, 2003). It is sometimes said that the AUC is a measure of prediction because a perfect AUC of 1.0 would indicate perfect discrimination between high- and low-risk individuals. While this is true, a very high AUC can be associated with a very low PPV when the base rate of the predicted event is low.
So if risk assessment is not about prediction, what is it about? Part of the answer comes from Pascal and Fermat. In their correspondence, they aimed to develop a method for the equitable sharing of a pot of potential winnings. When we assign a risk category to a group of people, we are in effect asking them to share risk as well. This is obviously the case for insurance and some forms of superannuation, but is also true in violence risk assessment where all the people in a high-risk category for violence have to experience the additional and potentially restrictive treatments so that the few among them who would have gone on to commit or experience harm can be treated (Large, Ryan, Nielssen and Hayes, 2008). When we assign persons to a low-risk category, each may be denied treatments which might be helpful in reducing some harmful actions or in other aspects of their life (Ryan, Nielssen, Paton and Large, 2010). To this extent, risk assessment is about the sharing and distribution of treatment resources and side effects rather than the prediction of particular events. In fact term ‘prediction’ cannot be found in the seminal works of Pascal, Fermat, and de Moivre – they understood that prediction was not possible. What they understood was that probability theory can assist in the equitable, even profitable, distributions of wins and losses.
Risk assessment and loss
Included in the earliest discussions of risk was the notion that the unit of risk is the unit of the loss rather than a probability expressed as a fraction. The monks of the Port-Royal Monastery used the example of lightning and suggested that lightning need not be feared because of its low probability. Today, we have the reverse problem where published accounts of risk are expressed almost entirely in terms of probability, correlation or AUC – all without reference to the losses involved. This might be partly attributed to Bernoulli’s notation of probability occurring between values of 0 and 1. It should be remembered that Bernoulli understood that that a probability is just a fraction, no more and no less. A fraction is meaningless unless it is a fraction of something. It is strange that this is so frequently forgotten in violence risk assessments that are defined solely as a probability. In reality, we are all familiar with the notion that risk has a monetary value: we all pay our insurance premiums in well defined amounts of hard-earned cash, and few people are aware of the probability that their car will be stolen, that their home will burn down or of the likelihood of any other loss they insure against.
The issue of loss in mental health risk assessment is rarely considered (Large and Nielssen, 2012). However, the crucial need to consider loss was well understood by both Bernoulli and de Moivre who stated explicitly that all the relevant losses and winnings ought to be assessed. While this is just common sense, and consideration of a complete set of risks is routine in many areas of medical decision-making, the diverse range of losses of differing magnitudes is a problem that has been ignored in violence risk assessment. In fact, even the term ‘violence risk assessment’ suggests that other risks, such as the risk of suicide, might not be considered. Furthermore, this comprehensive tally of probabilities and losses only makes sense if one course of action is balanced against another. A risk assessment should include not just the possibility of direct harms but also the unwanted consequences of any risk-based decision, variously referred to as side effects in medicine, externalities in economics or collateral damage in the language of the military. In risk assessment terms, externalities include the costs of detention and the infringements on the civil liberties of high-risk patients who would not have gone on to become violent.
When faced with a risk assessment, a clinician has to make something of a choice as to whether losses resulting from violence or self-harm should be the focus. Self-harm and violence do not have the same risk factors. For example, two recently published systematic reviews with meta-analysis cast some light on the risks faced by a patient who is admitted for treatment of a first episode of psychosis (Large and Nielssen, 2010; Large, Smith, Sharma, Nielssen and Singh, 2010). These reviews found that, while male sex, younger age and substance use were risk factors for violence, age and gender were not associated with suicide. Depressed mood was found to be protective against violence, but was (not surprisingly) a risk factor for suicide. Therefore, it is likely that a patient with an increased probability, relative to others, of self-harm also has a lower probability of violence relative to others and vice versa. Similarly, the risk factors for minor violence might differ from the risk factors for more severe violence (Large and Nielssen, 2011). Recently some research groups have started to consider the complexity of risk in mental health. For example, the ‘Short-Term Assessment of Risk and Treatability’ (Webster et al., 2006) addresses a wide range of possible harms. However, further research is needed to produce a risk assessment instrument to assess a variety of harms of differing levels of severity and with different base rates. Currently, actuarial risk assessment is unable, at present, to assess the possibility of more than one type of harm, or to consider the various risk factors associated with the varying levels of self-harm and harm to others. (Large, Ryan, et al., 2011).
Complicating matters further, although risk-assessment instruments might make a distinction between those at higher and lower probability of harm, they cannot necessarily meaningfully estimate the extent of the actual loss. There is enormous variation in both the intention and outcome of violent actions, from a minor push of a family member or fellow patient to a homicide. Even similar acts of violence can have very different outcomes: for example, a very minor assault to an able-bodied person might, in a vulnerable older person, result in a fall and eventual death.
Risk assessment and accuracy
A search of Medline reveals that the term ‘accuracy’ frequently occurs in publications about risk assessment. For example, a recent paper suggested that ‘imminent aggression in psychiatric hospitals may be able to be accurately predicted by psychiatric nurses’ (italics added) because ‘the prediction of any aggressive behavior irrespective of [the] type of aggression was significantly greater than chance (AUC=0.69)’ (Barry-Walsh, Daffern, Duncan and Ogloff, 2009). However, the Oxford English Dictionary defines ‘accurate’ as ‘precise, exact, correct’ – something violence risk assessments are not. This is not to say that accuracy has no place in probability theory. In fact, Bernoulli was concerned about just that. He wanted to discover the most accurate estimate of the underlying propensity to any particular outcome in a single trial. The use of large numbers of trials to estimate a natural propensity is the central issue in the law of large numbers and, by contrast, the results of one trial can never be informative. A watch can be accurate; a risk assessment for a future act of violence by an individual cannot be.
Risk assessment and certainty
Bernoulli and de Moivre were also concerned with the accuracy of their estimates of probability, but it was Bayes who asked how certain we can be, how much confidence we can have in our estimates, and it was Laplace who expressed Bayes’ theorem in its modern form. The surprising and somewhat controversial answer to this question of certainty is that it depends on our prior expectations. When we conduct a violence risk assessment, the positive predictive value of our high-risk categorizations depends on the base rate. The higher the base rate, the higher the positive predictive value and therefore the more certain we are in our risk predictions. At first glance, this seems objective enough, but do we really know the base rate of violence in any particular population and do we know whether a base rate estimate from an earlier study is, in Bayesian terms, a valid prior probability? For example, if we expect a person whom we are risk-assessing to be drawn from a population with a high rate of a particular type of violence, such as a simple assault, then Bayes tells us that we can be more certain about the accuracy of our assessments than if we were considering a rare form of violence such as homicide. This can be shown in the relationship between incidence and PPV (Szmukler, 2003). In essence, we can never have useful certainty about rare catastrophic events. This is a message that ought to have echoed down the centuries, after the killing of Jonathan Zito by Christopher Clunis, when violence risk assessment was introduced at the point of discharge from English psychiatric hospitals (Ritchie, 1994). ‘Stranger homicides’ by patients with schizophrenia, as was the case with Christopher Clunis, are so rare that they are all but impossible to research (Nielssen et al., 2011) – and the positive predictive value of any risk assessment for ‘stranger homicide’ means that as few as 1 in 35,000 risk predictions for stranger homicide will be associated with this particular harm (Large, Ryan et al., 2011).
Risk assessment and belief
A central controversy in statistics in the twentieth century was whether a probability is best thought of as a fraction or as a rational degree of certainty in belief. Bayes, Laplace and Keynes were more or less aligned in their view that probability is a measure of rational belief, whereas Fisher and his followers, sometimes referred to as Frequentists, were suspicious of the sort of prior assumptions made by Bayes. They saw probability as a fraction to be determined after a sufficient number of trials. While it has been argued that the heat has passed from this argument in statistics, it is worth considering two aspects of the disagreement between Bayesians and Frequentists in relation to violence risk assessment.
Consider the inverse probability problem of inferring the probability that a single patient will be violent from previously observed episodes of violence and non-violence. A Bayesian or Keynesian approach would infer that an individual would have a similar propensity to violence as the proportion of violent people in the population from which he or she was drawn. This is the underlying principle by which group characteristics are attributed to individuals. In one sense this is an obvious fallacy akin to assuming that all of the unseen marbles in a jar are pink because all of those sampled have been red or white. In contrast, a Frequentist would try to avoid making assumptions about unseen variables, working alone with what has been determined by the sampling procedure. While a Frequentist approach is less reliant on assumptions and it can lead to disregard of what is clearly useful and potentially causal information, Fisher – the chief originator of the Frequentist – rejected the suggestion that the association between smoking and lung cancer could be used as evidence that smoking was carcinogenic!
On the other hand, a strictly probabilistic risk assessment, based on objective risk factors and observations of large cohorts of similar patients as are considered in actuarial studies of risk can be seen as approximating a Frequentist position. While it is almost certainly true that this sort of assessment is statistically superior to clinical methods, the preference for clinical or actuarial risk assessment methods remains controversial (Buchanan, 2008). In reality, we probably cannot divorce prior beliefs from any risk assessment when faced with a particular patient. With respect to a single patient, we will observe a very small number of violent events, and there will always be doubts about how generalizable the results of any method of actuarial risk assessment are to a single person.
Arguably, even in an actuarial violence risk assessment, our degree of belief in the accuracy of the assessment ought to be tempered by our prior knowledge of the expected base rate. In Bayesian terms, even an actuarial risk assessment does not give us an answer as to the probability of a future harm because the base rates of violence vary and cannot be known with certainty until after the event, but serves to modify our prior belief in the likelihood for violence in the person being assessed.
Hence, we cannot really avoid our prior perceptions about the patient and must interpret the results of any actuarial risk assessment in this light. In fact, the arguable purpose of any risk assessment is to modify existing beliefs using new information, which is of course the essence of Bayesian thinking. Although a more strictly Frequentist approach to risk might be appropriate when very large amounts of data are available – for example, when recidivism by all released prisoners is considered – in relation to an individual patient, a Bayesian approach, using risk assessment to modify previously held opinion, has more practical use. To this extent, Bayesian approaches support structured clinical judgement, while Frequentist approaches support actuarial methods. A preference for a Bayesian view of probability in the face of complexity was neatly summarized by the well-known American statistician, John Tukey (1915–2000) who stated, ‘Far better an approximate answer to the right question, which is often vague, than an exact answer to the wrong question, which can always be made precise’ (Tukey, 1962).
Risk assessment and chaos
Leibniz’s correspondence with Bernoulli suggested that he believed that natural systems had a limited degree of predictability. By contrast, Laplace believed that all things were predictable. Few would now disagree with Leibniz or concur with Laplace. It is now well known that chaotic outcomes, those which deviate markedly from a previously observed mean in an unpredictable non-linear way, are fundamental in many forms of recursive systems operating with multiple feedback loops. Both Poincaré and Lorenz showed that tiny differences in initial conditions can make a large difference to observed outcomes. One implication of this for violence risk assessment is that increasing detailed assessments of the initial conditions might not necessarily lead to more accurate predictions. Furthermore, the most dreaded harms in mental health, homicide and suicide, are very rare. They are so rare that they can be considered to be outside the normal flow of events. Such unexpected and terrible events can be characterized as being non-computable. Chaos theory tells us that complex and recursive systems differ greatly in the extent to which they can be predicted. People are arguably among the most complex systems imaginable, where each person has billions of interconnected neurons in a much larger number of interacting neural circuits, which are influenced by sensory input, involving interactions with similarly complex other individuals. People are not predictable. Uncertainty in human behaviour is much more profound that the uncertainty in the roll of a die or the toss of a coin.
Conclusion
The limitations of risk assessment in mental health can be partially understood as resulting from the limitations of probability theory. Probability theory emerged in the course of considering issues of equity and not as an attempt to predict particular outcomes. Importantly, those included in a high-risk category will be likely to share resources and to suffer the side-effects of treatment. Further to this, risk assessments differ from simple assessments of probability in that they involve losses and gains. De Moivre and others go further, suggesting that a consideration of the probability of an isolated loss or an isolated gain in complex systems is not an assessment of the overall risk. Hence, a comprehensive assessment of risk must factor in all the important losses and probabilities – something existing risk assessment instruments do not attempt. Probability theory progressed from games of dice to situations in which the origin of the measured outcomes cannot be easily observed. This allowed the application of probability theory to human behaviour, the origins of which are obscure at best. However, where prior probabilities are not known, and can only be estimated, Bayes’ theorem tells us that prior expectations determine the confidence we can have in our risk judgements. The most important insight we can glean from this is that no matter how good our risk assessment is, the PPV or the degree of certainty we have in our high-risk categories will inevitably be determined by our knowledge of the base rate. To this extent then, risk assessment is as much about modifying our prior beliefs as it is about generating a numerical answer. Finally, probability theory, as it applies to complex systems such as people, has fundamental limitations that have been illustrated by Chaos theory. These limitations to our knowledge of future harms should be both acknowledged and respected.
Footnotes
Acknowledgements
I thank Dr Peter Arnold and Dr Christopher Ryan for their assistance with the manuscript.
Potential conflict of interest
Dr Large has received speaker’s fees from Astra-Zeneca to present his own research on the topic of risk assessment.
