Abstract
Background:
Artificial intelligence (AI) has the capacity to optimise the diagnosis and assessment of damage in inflammatory arthritis. The validation and performance metrics of AI algorithms should be interrogated in order to meaningfully map the existing data and identify evidence gaps.
Objectives:
To review the literature examining the use of AI algorithms to diagnose and assess damage in rheumatoid arthritis (RA) and psoriatic arthritis (PsA), using plain radiography (XR), ultrasound, computed tomography (CT) and magnetic resonance imaging (MRI).
Design:
A scoping review was performed for relevant English-language full-text articles or conference abstracts published prior to 12 September 2024 in the EMBASE, PubMed and Web of Science electronic databases.
Methods:
Abstract, full-text review and data extraction were completed by two authors independently utilising the Covidence platform. Data extracted included the study aims, publication type, study design, imaging modality, joints assessed, patient selection, diagnoses, algorithm(s) used, modifications and techniques utilised to improve model performance, testing and validation methodology and algorithm performance metrics. The review was completed in adherence to the Preferred Reporting Items for Systematic reviews and Meta-Analysis extension for Scoping Reviews (PRISMA-ScR) guidelines.
Results:
Two thousand and sixteen abstracts were screened following removal of 497 duplicates. Forty studies were included for data extraction, four of which were conference abstracts. A majority of studies were published in the last decade (88%) and a minority of studies involved patients with PsA (12.5%). Most studies utilised XR (n = 35), followed by MRI (n = 3), CT (n = 1) and USS (n = 1). Convolutional neural networks were the most commonly used algorithms. Only 15 studies (37.5%) used clearly described training, validation and testing datasets in algorithm development and testing.
Conclusion:
There is a paucity of data in the use of AI algorithms for the imaging assessment of PsA compared to RA. In RA and PsA, only a minority of published studies have utilised robust training, validation and testing methodologies in algorithm development. There was significant heterogeneity in the study populations, study design and performance metrics used, which limits comparisons between algorithms and the synthesis of results.
Trial registration:
Registration was not required in line with the PRISMA-ScR guidelines.
Introduction
The utility of imaging in inflammatory arthritis has evolved from plain radiography (XR) being used to understand the natural history of diseases and gain insights into disease pathophysiology, to the use of novel modalities such as magnetic resonance imaging (MRI) for diagnosis, prognostication and quantification of activity and damage. Accurate imaging evaluation in Rheumatoid Arthritis (RA) and psoriatic arthritis (PsA) is a cornerstone of high-quality clinical care and clinical trials, with significant implications in diagnosis, treatment decisions and prognosis. Imaging has the potential to facilitate early diagnosis by detecting subclinical or mild disease, allowing for earlier access to disease-modifying therapy. In diagnostic challenges, imaging may assist in differentiating RA from PsA and non-inflammatory musculoskeletal conditions, optimising therapeutic decision-making. From the perspective of prognostication, the presence of damage and damage progression is important to recognise as a predictor of future damage, long-term function and quality of life. In treatment decision-making, imaging not only allows for identifying patients at high risk of damage progression, but further allows for the detection of subclinical inflammatory disease to enhance clinical judgement. Finally, in clinical trials, structural damage endpoints are currently required for the robust demonstration of disease-modification. Centralised assessment utilising imaging endpoints with sensitive, specific and reliable scoring systems are important to support the validity of clinical outcomes in multicentre, multinational trials.
The imaging assessment of damage in peripheral inflammatory arthritis has historically utilised XR, a widely accessible and low-risk modality with the capacity to visualise relevant features such as joint space narrowing, erosions, osteopenia, osteoproliferation. Validated XR outcome measures such as the modified total sharp score (mTSS) for RA and the modified Sharp van der Heijde (mSvDH) score for PsA have formed the basis of damage assessment in research settings and are core outcomes in clinical trials. More recently, ultrasound (US) and MRI have gained traction; these modalities are able to provide a three-dimensional assessment of musculoskeletal structures and allow for the concurrent assessment of damage and disease activity. However, these imaging modalities and the scoring systems that are utilised are dependent on highly-trained readers, time-consuming and subject to inter and intra-reader reliability challenges. The potential trade-offs in cost-benefit, accessibility, sensitivity and discrimination are yet to be adequately established across research and clinical settings. 1
Irrespective of imaging modality or instrument utilised, the fundamental challenges of assessing damage include the limitations of human sensory perception and gold standard comparators, the time-consuming process of and variable access to training and the standardisation of imaging acquisition needed to ensure reliability and accuracy. In clinical settings, clinicians need significant exposure to imaging findings in healthy controls and degenerative diseases in order to be able to confidently utilise imaging in routine care without specialised radiology input. Advances in artificial intelligence (AI), and specifically deep learning, have the potential to revolutionise how we utilise imaging in routine care and research by improving the accessibility and performance metrics of such tasks.
Machine learning (ML) is a field of AI and is a unifying term for the use of computer algorithms to learn from data and generalise to unseen data. In musculoskeletal imaging, the most common ML tasks can broadly be categorised into supervised learning (e.g. classification and regression) and unsupervised learning (e.g. clustering). 2 Analogous to the choice of statistical test being dependent on the nature of the input data and desired outcome variable, the choice of ML algorithm is very much driven by the nature of the input data, the presence of labelling (supervised vs unsupervised) and the intended output, for example, a diagnostic label or a numerical estimation of joint damage severity. 3
Excellent primers have been published in recent years to aid clinicians in developing an understanding of ML algorithms that can be utilised in rheumatic diseases, the challenges that have stymied their widespread uptake and future directions for research.2–4 These reviews have provided broad overviews on how AI can be used within rheumatology and described general approaches to imaging assessment in rheumatic diseases.
The remit of this review is unique in that it focuses on RA and PsA and is an in-depth review of studies that have used ML for diagnosis, quantification of damage and assessment of damage progression, with inclusion of the joints in the hands and the feet. We specifically describe the algorithms used for specific tasks, joints assessed, modality utilised, patient numbers and selection, choice of comparator, the adequacy of algorithm validation and testing, strategies utilised to optimise the performance of algorithms and the performance metrics reported. In this review, we aim to highlight the challenges of interpreting existing ML studies for clinicians, such as the limited external validation of results, the heterogeneity in imaging modalities and the lack of standardisation in outcomes evaluated. We further discuss how learnings from previous research may be used to guide the design and reporting of future research, in order to overcome barriers to clinical implementation.
Methods
The scoping review was conducted in line with the Preferred Reporting Items for Systematic reviews and Meta-Analysis extension for Scoping Reviews (PRISMA-ScR) checklist. 5 As registration of scoping review protocols is not permitted on PROSPERO, a protocol was finalised prior to the literature search but not registered.
Search methods
The EMBASE, PubMed and Web of Science electronic databases were searched using search terms (Supplemental Material 1) as per protocol. These databases were selected to ensure comprehensive and methodologically robust coverage of relevant medical and computer sciences literature. Search results were imported into the Covidence review management platform with automatic removal of duplicates (n = 497, Figure 1). Additional articles were identified from the references (n = 4), and no grey literature was screened other than conference abstracts.

PRISMA search strategy.
Study selection
Eligibility criteria were as follows: (1) full text articles or conference abstracts, (2) English language, (3) published before 12 September 2024, (4) included a patient group with RA or PsA, (5) implemented a ML method, (6) utilised X-ray, computed tomography (CT), USS or MRI of the hands and/or feet, (7) aimed to diagnose or detect damage/damage progression.
The abstracts were independently reviewed by two authors from the search team for assessment of eligibility (A.A. and Y.M.). Full review of the short-listed conference abstracts and articles was conducted by two authors from the search team (A.A. and Y.M.). Inter-reviewer agreement was assessed by A.A., and there were no discrepancies requiring the input of a third author if required. Independent assessment of inter-reviewer agreement was not performed.
Data extraction
Two authors (A.A. and D.H.-M.) independently extracted the following data as per protocol: study aim, publication type, study design, imaging modality, joints assessed, patient selection, diagnoses, ML algorithm(s) used, modifications and techniques utilised to improve model performance, and testing and validation methodology and performance metrics of the algorithm at the final stage of algorithm development reported.
Reported performance metrics captured included but was not limited to: sensitivity, specificity, true and false positives and negatives, area under the receiver operator curve (AUC/AuROC), accuracy, precision, positive and negative predictive values, intra- and inter-rater reliability, correlation coefficients, root mean square error and root mean square deviation.
Data analysis
Extracted data were synthesised into tables summarising study characteristics and performance metrics, stratified by study aims. Data analysis was descriptive only. In accordance with the PRISMA-ScR guidelines, statistical summary measures were not calculated, and risk of bias assessment was not undertaken.
Results
The search strategy yielded 2016 articles for screening. Following removal of duplicates (n = 497), 1519 abstracts were screened. One hundred and sixty-seven full-text articles were screened for eligibility (Figure 1). In the final data extraction, 36 studies were included from the search strategy and 4 studies were identified from the references of the included studies and from content expert co-authors, yielding a total of 40 studies.
Of the included studies, 36 were full-text articles and 4 were abstracts. Thirty-five (88%) of these studies were published in the last decade and two-thirds (n = 27) were published in the last 5 years. A minority (n = 10) were published in rheumatology journals and four of these were conference abstracts. Most studies involved patients with RA (n = 38), two studies involved patients with PsA and three studies involved both PsA and RA patients.
The most frequently utilised imaging modality was XR (n = 35) followed by MRI (n = 3), high-resolution peripheral CT scans (HR-pQCT; n = 1) and USS (n = 1). No study combined or compared imaging modalities. The number of image sets utilised ranged from 15 to 12,866 across studies.
Study aims were categorised into: (i) diagnosis, (ii) detection of radiological abnormalities, (iii) scoring of radiological abnormalities, (iv) joint space measurement and (v) detection of joint damage progression (Table 2). Fifteen of the 40 identified studies (38%) had clearly defined training, validation and testing datasets (Tables 1 and 2).
Study design and methods.
ANN, artificial neural network; ASM, active shape model; CNN, convolutional neural networks; CSAE, convolutional supervised auto-encoder; DIP, Distal Interphalangeal; HR-pQCT, high-resolution peripheral CT scans; IP, (Interphalangeal) joints; MCPj, metacarpophalangeal joints; ML, machine learning; MRI, magnetic resonance imaging; PIP, Proximal Interphalangeal; PsA, psoriatic arthritis; RA, rheumatoid arthritis; SIFT, scale-invariant feature transform; SVR, support vector regression; XR, plain radiography.
Algorithm training, validation and performance.
AI, artificial intelligence; ASM, active shape model; AuROC, area under the receiver operator curve; CI, confidence interval; CNN, convolutional neural networks; FPN, feature pyramid network; HOG, histogram of gradients; HR-pQCT, high-resolution peripheral CT scans; ICC, intraclass correlation; IoU, intersection over union; MAE, mean absolute error; MCP, metacarpophalangeal; MIMO, multiple input multiple output; ML, machine learning; MRI, magnetic resonance imaging; MSGVF, multiple scale gradient vector flow; MTP, metatarsophalangeal; PsA, psoriatic arthritis; PsAMRIS, Psoriatic Arthritis Magnetic Resonance Imaging score; RA, Rheumatoid Arthritis; RAMRIS, Rheumatoid Arthritis Magnetic Resonance Imaging Score; RMSE, root mean square error; RMSQ, root mean square error; ROIs, regions of interest; SIFT, scale-invariant feature transform; SISO, single input single output; SNRA, seronegative RA; SPRA, seropositive RA; SVR, support vector regression; XR, plain radiography.
Diagnosis (n = 12)
Twelve studies investigated the use of ML for the diagnosis of inflammatory arthritis. Five of these studies included and adequately reported their training, validation and testing strategies (Table 2). Ten of the 12 studies included patients with RA, while two included patients with RA and PsA. A majority of the studies utilised XR (n = 8), followed by MRI (n = 2), USS (n = 1) and CT (n = 1) of the hands. All studies utilised control groups, including healthy controls, ‘no RA’, ‘osteoarthritis’ and ‘others’ (Table 2).
Ahalya et al. 25 included hand XRs of 100 patients with RA and 100 patients without RA in their study, which compared a number of custom and pre-trained Convolutional Neural Networks (CNNs), followed by the development of fusion modules of the best-performing CNNs with a scale-invariant feature transform algorithm and ML classifiers. The AuROC of the best-performing fusion model was 0.785.
Üreten et al. 31 included 1426 hand XRs of patients with RA, OA, Normal and ‘Other’, and used a YOLOv4 CNN to crop radiographs and a pre-trained VGG-16 CNN for two-, three- and four-way diagnostic classification. The AUC and accuracy for two-way classification (RA vs Normal) were 0.97% and 90.7%, respectively. The accuracy of four-way classification accuracy was 84.4%. Ma et al. (2024) included hand XRs of 9964 XRs (RA, OA and Normal) and used a pre-trained Efficient-Net-B0 CNN for classification, averaging predictions of a five-fold ensemble of models to achieve a three-way classification accuracy of 87.2%. 42
Fukae et al. 20 incorporated 1037 USS images from 239 patients with RA and ‘no RA’ with pathology results and clinical data to generate 2D arrays treated as image inputs into a pre-trained modification of the AlexNet CNN. The authors reported a precision of 91% and an agreement with three clinicians (Cohen’s Kappa) of between 0.79 and 0.87. The addition of clinical information did not improve the performance of the algorithm.
Two studies included patients with RA and PsA. One utilised 932 HR-pQCTs of the second metacarpophalangeal joint (MCPj) and included healthy controls and undifferentiated arthritis as comparators. This study utilised a convolutional supervised auto-encoder network for class prediction and included a training and 5-fold cross-validation dataset without a test set. 26 A second study by Folle et al. 27 included 649 MRI of the wrists, MCPjs and metatarsophalangeal joints in patients with seropositive (SPRA) and seronegative RA (SNRA), PsA and psoriasis. A ResNet CNN pre-trained on video understanding was trained on individual MRI sequences and fused into an ensemble model. An additional neural network trained on demographic and clinical features was studied to assess the impact of incorporating clinical information into a diagnostic algorithm. The authors reported AUROCs of the final models using various permutations of MRI sequences an inputs, achieving an AUROC of up to 74% for SNRA versus PsA, 75% for SPRA versus PsA and 67% for SNRA versus SPRA in the test dataset. The addition of clinical and demographic information did not improve algorithm performance.
Detection of radiological abnormality (n = 8)
Detection of erosions was assessed in four studies in RA. Two of these studies reported training, validation and testing strategies. Murakami et al. 16 included 159 hand XRs and used a Multiple Scale Gradient Vector Flow (MSGVF) snake algorithm to segment phalanges and a deep CNN trained on manually-set regions of interest (ROIs) for feature extraction and classification. On a test set of 30 XRs, the authors reported a true positive rate of 80.5% and a false positive rate of 0.84%. An average of 3.3 false positive ROIs were detected per case. Miyama et al. 29 included 40 hand and wrist RA XRs, using DeepLabCut for joint detection and a VGG16 CNN pre-trained on ImageNet and fine-tuned using joint images to detect erosions. Using orthopaedic assessment of erosions as ground truth, validation was performed on five patients. The authors compared models that classified joints independently versus models that utilised information from the same joint in the contralateral hand, the same joint type in the same hand (e.g. MCPjs) and the same joint type in both hands. The model utilising information from the same joint in the contralateral hand performed best (Precision Recall–AUC 0.73) by a small margin. The authors also evaluated joint space narrowing detection in the same study and found that the best-performing model incorporated information from the same joint class in the same hand. One other study assessed the presence of joint space narrowing in the hands and wrists of 915 RA patients. Wang et al. 32 utilised YOLOv3 and four CNNs for joint detection and EfficientNetB1 CNN for assessing the presence of joint space narrowing, purportedly using the mTSS to classify joints as normal, mild joint space narrowing or severe joint space narrowing. In a test set of 183 patients, an average accuracy of 91% was achieved, with accuracies being highest in the normal (0.90) and severe (0.91) joints. The accuracy was inferior (0.79) in ‘mild’ joint space narrowing, but the authors did not provide clear descriptions of these classes.
Three studies from one research group have assessed the presence of ankyloses and subluxation.14,35,39 These studies were conducted on XRs of the hands in RA patients and compared CNNs and Vision Transformer networks, demonstrating excellent and comparable performance metrics (Table 2), but utilised five-fold cross-validation for training and validation with no independent testing. While the validation dataset may have been used as an independent test set without any further tuning of the model, the published learning curve graph unusually demonstrated significant overfitting in model training without the expected associated drop in test accuracy.
No studies have investigated the presence of ankyloses, subluxation or osteoproliferation in PsA.
Scoring of radiological abnormality – Erosions (n = 13)
Eleven studies have reported on ML in the scoring of erosions in RA using XR, while one used XR in PsA patients and one used MRI in PsA and RA patients. Six studies adequately reported training, validation and testing.
In RA, three studies used XRs of the hands. Hirano et al. 17 utilised 216 radiographs and implemented a cascade classifier using Haar-like features to detect joints and a CNN to assign radiographic damage scores, with the mTSS erosion score serving as ground truth. They reported 70.6%–74.1% exact agreement and 84.3% close agreement (score ±1) between the model and two readers, and correlation (r) of 0.54–0.75. The agreement between model and individual readers was comparable to the agreement between readers. Rohrbach et al. 18 implemented a modified VGG16 CNN pre-trained on ImageNet to estimate the Ratingen erosion score and similarly achieved agreement between their model and reader scores that was comparable to the agreement between readers (Table 2). The authors found that a weighted categorical cross-entropy model outperformed a non-weighted model and that in the context of having a very large training dataset, transfer learning did not improve the performance of their model. Radke et al. 36 utilised CNNs (RetinaNet with a ResNet and a feature pyramid network) to estimate the mTSS erosion score, and assessed the impact of adaptive intersection over union (IoU) thresholding during training on accuracy. The confusion matrices were similar between the model and the readers as between two readers (Table 2). The agreement between model and rheumatologist was highest in joints that were normal (score 0, agreement 0.99) and significantly eroded (score 5, agreement 0.96; Table 2). The use of adaptive IoU thresholding was found to improve model accuracy.
One study utilised XRs of the hands and feet. Sun et al. 30 reported the results of a large crowd-sourcing AI challenge (RA-DREAM) for radiographic damage scoring of the hands and feet that included a post-challenge independent validation. With 13 submissions in the final challenge round, most of which involved the use of CNNs, a submission that utilised ResNeT CNNs enhanced by XG boosting achieved a weighed root mean square error (RMSE) of 0.43 for the estimation of the erosion score.
In PsA, Janiczek et al. 40 utilised XR of the hands and feet and implemented YOLO-V8 for joint detection and DenseNet 122 CNNs to achieve an MSE of 0.38 ± 2.01 at the level of the joint, with an intraclass correlation (ICC; (95% CI) of 0.85 (0.93–0.95); this was comparable to inter-reader ICCs.
A further solitary study has reported on the use of MRIs of the hands and wrists of patients with RA and PsA, using the Rheumatoid Arthritis Magnetic Resonance Imaging Score (RAMRIS) and Psoriatic Arthritis Magnetic Resonance Imaging score (PsAMRIS) erosion scores as ground truth. Schlereth et al. 44 utilised 431 MRIs from 187 patients in their ResNet-3D models, which were pre-trained on a video classification task, for score prediction. In the independent validation, the authors reported a macro AUC of 87.89% and a balanced accuracy of 43.56%. The macro-AUC, or ‘arithmetic mean of class-wise AUCs in multi-label learning’, 47 and balanced accuracy (balanced for baseline imbalances in class numbers) were better in the reader exercise compared to the model. The authors found that increasing the training dataset size improved model performance, and that using non-contrast sequences generated similar results to contrast-enhanced sequences. The implementation of a SwinTransformer architecture pre-trained on video classification tasks48,49 was also tested, but within the limitations of the small training set, no improved performance was seen.
Scoring of radiological abnormality – Joint space narrowing (n = 10)
A total of 10 studies were identified, with all studies utilising XRs. In RA, nine studies have assessed JSN, of which four adequately reported training, testing and validation. Three of these studies assessed XRs of the hands and one assessed XRs of the hands and feet.
Hirano et al. 17 reported a percentage of exact agreement between readers and their model of 49.3%–65.4% for joint space narrowing, with a percentage of close agreement (score ±1) of 64.0%–85.3%, and a correlation (Pearsons r) of 0.72–0.88; these results were comparable to the model’s performance for erosion score estimation. Deimel et al. 45 conducted a large study (5191 XRs) assessing erosions at the MCPs and PIPjs of the hands on XRs using a SpatialConfiguration-Net CNN for joint localisation and modifications of the VGG16 and DenseNet CNNs for scoring. The achieved mean accuracy was 80.5% at the metacarpophalangeal joints and 72.3% at the proximal interphalangeal joints, and a score difference of >1 was present in <2% of joints, however, no information regarding severity and distribution of damage severity in the cohort was available in the abstract. Wang et al. 32 implemented YOLOv3 and 4 CNNs for joint detection and EfficientNetB1 CNN for scoring of erosions in a study of 915 XRs. Used a modified version of the mTSS in which scores 1–2 and scores 3–4 were grouped together as a comparator, the authors reported a weighted average precision of 0.88.
In the RA-DREAM crowd-sourcing experiment reported by Sun et al., 30 the winning team in the joint space narrowing score sub-challenge implemented a U-net model for segmentation and a random forest model for score estimation, achieving a weighted RMSE of 0.38.
Janiczek et al. 40 assessed joint space narrowing scores in the hands of PsA patients, achieving an excellent ICC (95% CI) of 0.96 (0.96–0.96) with readers and a mean square error of 0.14 ± 0.64 at the joint level.
Scoring of radiological abnormality – Overall damage (n = 4)
Damage as a composite outcome has been assessed in four RA XR studies, three of which adequately reported training, testing and validation dataset.
Wang et al. 33 conducted a study using 1618 hand XRs of patients with RA or ‘suspected RA’, comparing 3 CNNs and its modifications and assessing the impact of transfer learning on model performance. With the mTSS serving as ground truth, the best-performing model (ResNet-DWise50) achieved a mean absolute error (MAE) of 14.90 and an RMSE of 22.01; both metrics improved with the implementation of transfer learning. Bo et al. 19 aimed to validate these findings using models (ResNet-50, MobileNetV2 and ResNet-34) pre-trained on the RNSA Paediatric Bone Age Challenge dataset and ensembled models, further assessing smooth loss and MSE loss of function. Using 3818 hand XRs of patients with confirmed RA, their best performing model (ensemble ResNet-50: FBs-1, ResNet-34:finetuned and MobileNetV2:IRBs-3 finetuned with MSE loss) achieved an MAE of 12.57 and an RMSE of 18.02 in their test cohort, which was on par with an experienced reader. Utilising hand and feet XRs of RA patients in the RA-DREAM 30 crowd-sourcing challenge, the best performing model utilised a DenseNet201 CNN and achieved a weighted RMSE of 0.44.
Damage as a composite outcome in PsA was also reported by Janiczek et al. 40 in PsA. The ICC with readers at a patient level was 0.92 (95% CI: 0.92–0.93), with slightly better correlation in the hands (ICC 0.93 (95% CI: 0.92–0.93)) compared to the feet (ICC 0.88 (95% CI: 0.87–0.89)). The model versus reader ICCs were comparable to inter-reader ICCs, with smaller mean square errors (Table 2).
Joint space measurement (n = 3)
Three studies have assessed the measurement of joint space, two in RA and one in PsA. Two of these studies utilised shape models in their algorithm9,24 and two utilised neural networks.6,24 None of these studies utilised training, testing and validation datasets (Table 2).
Scoring of radiological damage progression (n = 2)
The assessment of radiological progression has been studied in two RA studies. Both studies assessed progression in joint space narrowing only and only one study had a complete training, testing and validations reported. Wang et al. 37 utilised a supervised U-net++ network for joint segmentation and a ResNet-like network for image registration and quantification of joint space change in hand XRs. Using the Genant-modified Sharp Score (GSS) as ground truth, the authors validated their model in 30 patients (15.2% of joints with radiographic progression) and demonstrated a small but statistically significant correlation between the model’s estimate and the change in GSS (r = 0.22, p = 0.0003).
Discussion
ML learning techniques hold the promise of improved accuracy, efficiency and reliability in diagnosis and damage quantification of inflammatory arthropathies. However, the utilisation of ML in rheumatology imaging has been largely limited to research settings. This scoping review provides an overview of the evidence to date in RA and PsA, summarises the performance metrics achieved for individual tasks and interrogates the methodology used in algorithm training, validation and testing.
A variety of ML techniques have been utilised in the literature, and the choice of technique is dependent on the nature, quality, labelling and quantity of the input data, the aim of the algorithm, the output (classification, regression) and the preferred techniques of the modeller. The most commonly used technique identified was artificial neural networks (ANNs), in which layer(s) of neurons are used to connect an input layer (the data we have) to an output layer (the outcome of interest). The earliest studies identified in our scoping review utilised ANNs. 6 The classical architecture of a neural network is one in which every neuron in each layer connects to every neuron in the subsequent layer, while recurrent or residual networks refer to architectures in which the neurons are linked in different ways. 4 Neural networks that have more than two layers (i.e. input layer, output layer and typically multiple hidden layers) are referred to as deep neural networks.
The use of neural networks in the interpretation of imaging data was revolutionised by the advent of CNNs and the utilisation of more complex and deeper neural architectures such as AlexNet. 50 CNNs were the most commonly used neural network in this scoping review, and were used in 29 of the 40 identified studies. CNNs are often used for classification and can classify input image data into a diagnostic category (e.g. PsA or RA or normal) or a specific damage score at the level of the joint or the patient level. Detector/segmentation CNNs are also frequently used to apply bounding boxes over joint regions or to segment ROIs, prior to a second CNN or alternative ML technique being used in ensemble. The key feature of a CNN is the convolutional kernel, which is an array of weights that interact with the pixel intensities of an input image by moving across it systematically in order to detect features in the image, such as edges. 2 In deeper layers further along the network, the kernels convolve with the feature maps generated by previous convolutional layers. In deep learning, the weights in the kernels, and thus the features they detect are learnt as the model is trained. In CNNs, the convolutional kernels in earlier layers are thought to be trained to detect more lower level image features, while the deeper layers are thought to detect higher level image features. 2
The majority of the studies identified in this literature review focused on diagnosing RA, predominantly with plain radiography; all utilised CNNs. The performance metrics achieved in individual papers are summarised in Table 2. It is however inadvisable to make conclusions regarding the superiority of one algorithm’s performance over another given the significant heterogeneity in and the inadequate reporting of baseline characteristics of the study population (e.g. severity of imaging abnormalities, disease stage, comorbid musculoskeletal findings, etc.) and the nature of the comparator group (e.g. osteoarthritis vs healthy controls vs ‘no RA’). For example, an algorithm trained and tested on a dataset with patients with established RA with significant damage may have a superior diagnostic performance than a dataset with early RA and no damage. Furthermore, some studies undertook two-way classification (e.g. RA vs normal) while others undertook three-way classification (e.g. RA vs PsA vs normal). Typically, studies with two-way classifications report superior performance metrics than studies with higher number of classification groups. This is important to note for real-world implementation where there may be a much higher number of diagnostic groups and where patients are more likely to have co-existing radiographic changes (e.g. RA and OA). Finally, the performance metric of choice (e.g. AuROC, accuracy, precision, etc.) varies between studies, and it can be difficult to be sure that the performance metric selected for reporting has not been cherry-picked. In synthesising the available data, it appears possible to achieve accuracy and precision of >80% and a AuROC of >0.785 in the diagnosis of RA using plain radiography (Table 2). In one study utilising US 20 for RA diagnosis, the addition of clinical information did not appear to improve the performance of the model.
No studies have examined the performance of a ML algorithm to diagnose PsA using XRs or US. Two studies assessed the diagnosis of PsA in other modalities; both included patients with RA and were carried out by the same centre. One study utilised high-resolution pQCT of the second MCP and reported promising results but did not have independent validation and testing datasets. 51 The second study utilised MRI and achieved an AuROC of 74%–75% for differentiating SPRA or SNRA from PsA. 27 Importantly, the authors also found that adding clinical data to the model did not improve its performance, and the removal of contrast-enhanced sequences did not result in a significant loss in model performance.
This scoping review elicited very few studies with complete training, validation and testing datasets for the detection of erosions16,29 and joint space narrowing,29,32 respectively in RA. A notable finding in one study of 40 patients that included both erosion and joint space narrowing detection was that incorporating information from other joints could improve the performance of an algorithm, which is something readers do intuitively when they assess radiographs. 29 This hypothesis is also supported by the findings of Li and Guan 35 who looked at the scoring of erosions in hand and feet XRs of patients with RA. The study reported training with a tenfold cross-validation approach; the model utilised won the RA-DREAM crowd-sourcing challenge for the joint space narrowing sub-challenge (Table 2).
Lesion detection studies largely utilised CNNs, however some utilised a shape model for segmentation along with CNNs.16,24 Shape models are statistical models of the range of possibilities for the shape of an object (e.g. a phalanx). They can be utilised to segment a ROI or to detect variations from ‘normal’ or baseline. Shape models utilised in this scoping review include the MSGVF snake algorithm, 16 Active Shape Model-driven snake algorithms 7 and an active latent space shape model that utilised a Bayesian Gaussian process latent variable model alongside U-Net. 24 However, segmentation can also be performed by CNNs; a majority of the finalists in the RA-DREAM crowd-sourcing challenge utilised a CNN for segmentation prior to score estimation.
Conventional imaging scoring systems such as the van der Heijde modifications of the Sharp score in RA were developed in order to semi-objectively quantify the severity and progression of erosions and joint space narrowing on XR. The RAMRIS and the PsAMRIS are used in MRIs to assess inflammatory activity, bone erosions and proliferation as a binary outcome in PsA. These instruments have been used in clinical trials to score damage and assess damage progression over time, and have served as the gold standard for ML algorithm validation. In interpreting the data regarding the assessment of damage and its progression however, it is important to bear in mind the context in which these instruments were developed. The features assessed and the joint regions selected for scoring in these instruments were selected to ensure they could be scored accurately and reliably (inter-rater and interrater) by human readers. For example, many wrist joints are either combined or excluded in the modified Sharp scores due to the overlapping nature of the carpal bones. The features selected also have to progress enough to be detectable by the human eye over time in order to be worthwhile measuring in the context of a clinical trial, assuming they can be measured accurately and reliably. For example, new bone formation is not routinely assessed on XR in clinical trials of PsA despite it being an important radiological feature of the disease because it was not found to change significantly in phase III RCTs in PsA when assessed by readers. 1 Therefore, while these instruments are certainly the best available current gold standard, they may not be the best representation of ground truth for use in the context of ML validation. The progress of ML may ultimately require alternative reference standards rather than human-derived scores. Techniques such as annotation of bones to create statistical shape models or to train CNNs for example, may provide a more valid comparator for ML validation. This may allow for improved sensitivity in the detection of pixel-level abnormalities (i.e. changes in erosion, joint space area or volume over time), a hypothesis that warrants further investigation.
In the scoring of erosions and joint space narrowing in RA using plain radiography, a variety of ML algorithms have been utilised including a support vector machine using histogram of gradients and support vector regression with histogram of gradients, a cascade classifier using Haar-like features, a tree-based model and a range of CNNs. There was notably significant variability in the choice of metrics used to assess algorithm performance (Table 2). The RA-DREAM crowd-sourcing challenge was the most rigorous of these studies, utilising radiographs of the hands and feet and including an external independent validation cohort. 30 The winning algorithms were able to achieve low errors, with a RMSE of 0.43 for erosion scores (ResNets for segmentation and classification), a RMSE of 0.38 for joint space narrowing scores (U-net for segmentation and CNN and ensemble forest models for scoring) and a RMSE of 0.44 for overall damage score (DenseNet 201 CNN). Importantly, the authors found that ensembling the top 9 performing models systematically outperformed the top performing model and moderate concordance indices were achieved in the post-challenge independent validation.
Other factors that appear to improve model performance when the utilising ML for RA damage scoring may include the use of an adaptive ‘intersection over union’ threshold during training 36 and the use of transfer learning, particularly in smaller datasets. 33 Commonly used datasets for transfer learning included ImageNet, 52 the Paediatric Radiological Society of North America Bone Age dataset 53 and CheXpert. 54
In studies assessing damage which reported their performance metrics in a confusion matrix, a further observation was that the agreement between reader and algorithm was typically high when damage was minimal or severe, but lower in the intervening stages of damage. 36 The overall performance metric therefore can be significantly impacted by the distribution of damage severity, or the class balance, in the cohort. There is typically also a significantly higher agreement when the performance metric used is percentage of close agreement in score rather than percentage of exact agreement in score.17,18
Only one study has examined the detection of damage in XR in PsA. Janiczek et al. 40 utilised CNNs (YOLO-V8 for joint detection and DenseNet 122 for classification) on XRs of the hands and feet from PsA RCTs, achieving comparable ICCs to readers (ICC (95% CI) 0.85 (0.83–0.85)) with a lower MSE ± SD than readers (0.38 ± 0.2 vs 0.64 ± 2.49). 40 These results are comparable to the inter-reader and intra-reader ICCs reported in phase III PsA RCTs. 1 Schlereth et al. 44 included both patients with RA and PsA and utilised MRIs of the hands and feet and ResNet-3D models pre-trained on video classification in MRI. They achieved a macroAUC (three-way classification) of 87.89% that was comparable with readers, however, the balanced accuracy was superior amongst readers. The authors also found that using a vision transformer did not improve model performance in the context of their moderate study size.
The prevention of damage progression is an important management priority for patients with inflammatory arthritis given the associations between damage and physical function. 55 It is a key outcome for the regulatory approval of disease-modifying anti-rheumatic drugs, in clinical trials comparing therapeutic options, and in routine clinical care. Wang et al. 37 developed an algorithm with a supervised U-net model and incorporating a ResNet-like network for image registration and subsequent quantification of joint space change, but found a low correlation between their algorithm and change in the Genant-modified Sharp Score (r = 0.2167, p = 0.0003). More recently, Venäläinen et al. 56 have reported success in detecting change in damage scores in hand and feet XRs of patients with RA. The authors modified their model (an entrant to the RA-DREAM challenge) using YOLOv3 for joint detection, a gradient boosting machine for joint labelling and DenseNet 121 and DenseNet 169 CNNs for joint prediction. In a cohort of 57 RA patients with an average interval of 4.6 years between XRs and an average mTSS increase by 5.3, the authors reported a Pearson’s r of 0.74 (95% CI: 0.59–0.84, p < 0.001). However, significant post-processing strategies were required to the model outputs, including imputation and multiplication with a scaling factor, in order to avoid the systematic underestimation of joint scores.
The assessment of radiographic change is typically performed less reliably by readers than the assessment of radiographic scores at a single timepoint, making it a key area in which ML algorithms could conceivably outperform readers. 1 The scarcity of ML studies assessing radiographic change over time is a major evidence gap in the literature, as is the paucity of data in PsA. Patients with PsA were included in 5 of the 40 identified studies: 2 related to PsA diagnosis and also included patients with RA,26,27 1 evaluated the measurement of joint space in PsA patients with normal radiographs 24 and 2 assessed damage scores using XR 40 and MRI 44 of the hands and feet.
A majority of the studies identified in our literature review (62%) did not have adequately-described training, validation and testing strategies, which is a key methodological limitation in the field. Assessing for the presence or absence of these partitions of data and how they are utilised is a critical part of appraising AI literature. In ML, the modeller typically chooses their algorithm(s), which may be an existing algorithm, a modification of an existing algorithm or a custom-built algorithm. They then determine the hyperparameters for the model, which are set prior to the training of the model and can be updated during the runtime of the model. In CNNs, these include the number of layers and neurons, initialisation, drop-out rates, learning rates, batch sizes, epochs and optimisation algorithms. The model is then trained on a set of data known as the training data, which uses an optimisation algorithm (e.g. Adam Optimizer) to generate the models’ parameters to minimise loss. In neural networks, parameters include the weights and the biases of the neurons. 2 Once an algorithm is trained, it is validated on a validation dataset; at this stage the hyperparameters may be tuned further to avoid overfitting or underfitting. Finally, the data are tested on an unseen dataset, which should ideally be large and representative of the population of interest. In this testing stage, no further alterations are made to the parameters or hyperparameters, and therefore the results are representative of the model’s performance. In best practice, further testing is conducted on an external dataset or a holdout set to confirm the findings. Only one study, RA-DREAM, utilised an independent external dataset. 30 The datasets utilised in this study have the potential to revolutionise the field by providing a high-quality large dataset for standardised external validation if they were accessible to researchers more widely, and allow for direct comparisons across models. Similarly, large imaging datasets from PsA pharmaceutical clinical trials could be repurposed in a similar way.
The majority of studies typically only had two datasets, one referred to as a training dataset and one variably referred to as either a testing or validation dataset. While it may be reasonable to only have a training and validation dataset, this would only be appropriate in the event that no further changes were made to the model at the validation stage and that the validation dataset was independent. The lack of clarity in these descriptions makes it challenging to know whether the reported results have selectively reported from a range of testing runs, and a well-performance algorithm in testing may fail in the validation stage if there has been overfitting in the testing phase. Efforts to standardise the reporting of AI studies in medical imaging have progressed significantly in large part due to efforts by the Checklist for Artificial Intelligence in Medical Imaging (CLAIM) group, 57 and will result in continuous improvement in the quality of the literature in years to come if mandatory implementation is required for publication.
A key challenge in the AI is the lack of explainability in deep learning models, which has flow on impacts on its regulatory approval and deployment in clinical practice. 58 In medical imaging, a commonly used approach is a saliency map. Saliency maps highlight the region(s) of the images that were most relevant for the model to generate its output and provides a measure of how relevant the regions are. Commonly used methods include occlusion- or perturbation-based methods (e.g. SHaplep Additive exPlanations, or SHAP) and gradient-based methods (e.g. Gradient-weighted Class Activation Map, or GradCAM). For example, Folle et al. 26 in their study of HR-pQCT of the second MCPj of patients with RA, PsA, healthy controls and undifferentiated arthritis, implemented the GradCAM to demonstrate that ‘erosions in the bare area and osteoproliferative changes in the ligament/capsule insertion sites’ were critical for disease classification. 26 Bo et al. 19 also utilised GradCAM in RA, finding false positives in which highlighted joints did not demonstrate pathology and missed pathology in the wrists.
Ultimately, such approaches do rely on the interpretation of the human assessing the saliency map and is potentially fallible in terms of confirmation bias. An alternative, albeit currently experimental approach, may be to implement natural language explanation models. This could conceivably allow the model to explain the findings that contributed to its decision-making, as a radiologist might.(42). Ultimately findings generated by approaches to address explainability in clinical decision-support systems need to be comprehensively reported and cautiously interpreted. 59
In this paper, we aimed to present an updated literature review with all available data in order to highlight the advances in the field and generate recommendations for moving the field forwards. Our key finding is that the majority of studies in the field are methodologically inadequate given the absence of clearly defined training, validation and testing datasets, and questionable validation practices. While ML-based scoring approaches often achieve reader-level agreement, particularly at extremes of disease severity, the variability in ground truth definitions, scoring systems and reported performance metrics makes cross-study comparison challenging. The use of extensive post-processing techniques has shown promise in improving the performance of algorithms, however in the absence of adequate external validation, the possibility of overfitting and loss of generalisability cannot be excluded. We further emphasise the relative dearth of data in PsA and non-XR modalities, and the observation that the addition of clinical data may not improve the performance of ML algorithms. Further research is required in PsA before ML-based imaging tools can be utilised outside research settings.
We further recommend that mandatory adherence to reporting guidelines for such as the CLAIM checklist in ML publications in order to accelerate progress in the field and minimise the publication of inadequately validated studies. Accessible heterogeneous longitudinal validation datasets and datasets which do not exclude co-morbid conditions such as osteoarthritis would improve the clinical relevance of ML algorithms. Furthermore, very little can be concluded regarding the use of non-XR modalities as only 12.5% of the studies reviewed utilised USS, MRI or CT. The reasons for this may include accessibility to equipment, funding and operator and/or reader training, and the relative recent development of scoring instruments for other modalities. Future research should incorporate these advanced modalities where feasible in order to gain insights into the relative diagnostic and prognostic yields.
This review has several limitations. Firstly, it was a scoping review rather than a systematic literature review. A scoping review was chosen due to the heterogeneous nature of the data and the anticipated challenges in data synthesis, but is limited by the lack of risk of bias assessment. As such, our findings should be interpreted as descriptive, rather than an evaluation of algorithm superiority. A systematic review has recently been published examining the use of ML more broadly in rheumatic musculoskeletal diseases – the focus of the review differed from this publication in that it was limited to hand imaging and full-text publications. While the authors had a different approach to the data synthesis, they similarly highlighted the paucity of external validation in identified studies without utilising a formal risk of bias assessment tool. 60
Additionally, we excluded studies that assessed the axial joints and other peripheral joints; notably the preponderance of evidence for the assessments of joints such as knees and hips is in osteoarthritis. We acknowledge that the exclusion of axial imaging is a major limitation, and this would be an important piece of work once definitions for axial PsA have been validated.
Many challenges remain in making progress in the field of imaging for rheumatological diagnosis and damage assessment. The need for standardised reporting of ML studies is crucial and clinicians face the challenging task of essentially learning a new language in order to appraise the literature. Large volumes of high-quality accessible training data are imperative in order to improve algorithm performance as well as to serve as an external validation dataset that could be used to draw fair comparisons between algorithms, however regulatory barriers must first be addressed. Finally, the importance of close and enduring collaboration between ML scientists and clinicians is imperative across all stages of algorithm development in order to facilitate transitioning of these algorithms into real-world applications.
Conclusion
ML has the potential to revolutionise the diagnosis and damage assessment in inflammatory arthritis from the perspectives of accuracy, sensitivity, reliability and cost. While many relevant studies were identified in this literature review, particularly for RA, there was significant heterogeneity in study methodology, patient populations and performance metrics reported. Limited data were identified for the assessment of radiographic damage and in the assessment of PsA in general. As the evaluation of ML algorithms in rheumatology continues to expand, it is important for clinicians to be aware of the pitfalls in the interpretation of study results. International collaboration in developing large diverse labelled longitudinal imaging datasets for standardised external validation, consensus adoption of standardised reporting checklists, an expansion of PsA-focused studies and broader utilisation of imaging modalities should be a focus of researchers in the field.
Supplemental Material
sj-docx-1-tab-10.1177_1759720X261460122 – Supplemental material for Artificial intelligence in the diagnosis and assessment of damage in peripheral inflammatory arthritis: a scoping review
Supplemental material, sj-docx-1-tab-10.1177_1759720X261460122 for Artificial intelligence in the diagnosis and assessment of damage in peripheral inflammatory arthritis: a scoping review by Anna Antony, Dylan Henley-Marshall, Yi Mon, Neill Campbell, Tony Shardlow, Bartlomiej W. Papiez and William Tillett in Therapeutic Advances in Musculoskeletal Disease
Supplemental Material
sj-docx-2-tab-10.1177_1759720X261460122 – Supplemental material for Artificial intelligence in the diagnosis and assessment of damage in peripheral inflammatory arthritis: a scoping review
Supplemental material, sj-docx-2-tab-10.1177_1759720X261460122 for Artificial intelligence in the diagnosis and assessment of damage in peripheral inflammatory arthritis: a scoping review by Anna Antony, Dylan Henley-Marshall, Yi Mon, Neill Campbell, Tony Shardlow, Bartlomiej W. Papiez and William Tillett in Therapeutic Advances in Musculoskeletal Disease
Footnotes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
