Abstract

To the Editor,
Ray, Dark and Kisely offer a thoughtful account of the pressures reshaping our speciality, and we share their conviction that psychiatry must redefine its value around what cannot be automated (Ray et al., 2026). One claim, however, deserves challenge, because it concedes ground that the mathematics says was never contestable. The authors grant artificial intelligence the tasks of risk prediction and risk flagging, suggesting that here the machine will be ‘faster, more consistent, and less biased’ than the clinician. We think this misreads where the difficulty lies.
The predictive value of any test for a rare outcome is governed not by the accuracy of the instrument but by the prevalence of the outcome. When the base rate is low, false positives swamp true positives, and the positive predictive value collapses however sensitive and specific the test is. Maguire and Looi (2025) have recently set out this reasoning for diagnostic testing in this region’s literature, showing that a 99%-sensitive, 99%-specific test for a condition of 0.1% prevalence yields a positive predictive value of only 9%. We demonstrated the same constraint empirically for suicide risk scales, where positive predictive values ranged from 0.3% to 1.8% in low-prevalence populations and the number needed to intervene reached 325 (Gale et al., 2019).
The decisive point for the authors’ argument is that this floor does not move when the classifier improves. A systematic review and meta-analysis of 53 machine learning studies, covering some 35 million records, has now found that these algorithms misclassify more than half of those who later die by suicide or re-present for self-harm as low risk, and that a high-risk flag carries a positive predictive value of around 6–17%; the authors concluded the algorithms were no better than the traditional scales (Spittal et al., 2025). A closed-form theorem, a meta-analysis of rating scales, and a meta-analysis of machine learning thus converge on a single answer across a generational advance in the technology. Notably, those algorithms achieved areas under the curve of 0.69–0.93 – discrimination that looks more than respectable – while failing at the only task that matters at the bedside, because the global metric is prevalence-independent and the patient is not. The claim that an algorithm is ‘less biased’ is, in this setting, close to incoherent: the collapse in predictive value is the prior reasserting itself through the likelihood, irrespective of any fairness internal to the model.
This strengthens rather than weakens the authors’ case, but relocates it. If risk flagging is Bayes-limited regardless of who or what performs it, then it was never the defensible psychiatric act. The defensible act is the one that precedes any test: setting the prior by selecting whom it is even coherent to assess – the clinical stratification that distinguishes the high-prevalence group in which a test may inform from the unselected population in which it cannot (Gale et al., 2019; Maguire and Looi, 2025) – and then signing the proportionate decision under the irreducible uncertainty the arithmetic guarantees will remain. That judgement, and the accountability that attaches to a named clinician who must defend it, is precisely what neither an algorithm nor a deferred appeal to the ‘circle of concern’ can supply. Psychiatry’s future in the management of risk lies not in faster prediction but in owning the decision that prediction cannot make – a judgement that can only be owned within a clinical relationship, not outsourced to a classifier.
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The authors received no financial support for the research, authorship and/or publication of this article.
