Abstract
Objectives:
The presence of nodal metastases in patients with papillary thyroid carcinoma (PTC) has both staging and treatment implications. However, lymph nodes are often not removed during thyroidectomy. Prior work has demonstrated the capability of artificial intelligence (AI) to predict the presence of nodal metastases in PTC based on the primary tumor histopathology alone. This study aimed to replicate these results with multi-institutional data.
Methods:
Cases of conventional PTC were identified from the records of 2 large academic institutions. Only patients with complete pathology data, including at least 3 sampled lymph nodes, were included in the study. Tumors were designated “positive” if they had at least 5 positive lymph node metastases. First, algorithms were trained separately on each institution’s data and tested independently on the other institution’s data. Then, the data sets were combined and new algorithms were developed and tested. The primary tumors were randomized into 2 groups, one to train the algorithm and another to test it. A low level of supervision was used to train the algorithm. Board-certified pathologists annotated the slides. HALO-AI convolutional neural network and image software was used to perform training and testing. Receiver operator characteristic curves and the Youden J statistic were used for primary analysis.
Results:
There were 420 cases used in analyses, 45% of which were negative. The best performing single institution algorithm had an area under the curve (AUC) of 0.64 with a sensitivity and specificity of 65% and 61% respectively, when tested on the other institution’s data. The best performing combined institution algorithm had an AUC of 0.84 with a sensitivity and specificity of 68% and 91% respectively.
Conclusion:
A convolutional neural network can produce an accurate and robust algorithm that is capable of predicting nodal metastases from primary PTC histopathology alone even in the setting of multi-institutional data.
Keywords
Introduction
Papillary thyroid carcinoma (PTC) is the most common thyroid malignancy, comprising 80% to 85% of all thyroid cancers. 1 Population data from the NIH demonstrated an incidence of 13.7 per 100 000 persons in 2019. 2 Surgery is the recommended initial treatment for most PTC. The presence and extent of lymph node metastases in PTC is an important factor in postoperative management, particularly in regards to decision making surrounding radioactive iodine (RAI) ablation. Guidelines from the American Thyroid Association (ATA) cite a 20% risk of structural disease recurrence with 5 or more lymph node metastases. 3 Despite the prognostic utility of this information, lymph nodes are not routinely removed during thyroid surgery.
Preoperative imaging has low sensitivity when identifying metastatic lymphadenopathy, particularly in the central neck.4 -8 Surgical removal of lymph nodes is the only other method to identify metastatic adenopathy, but the downsides of elective central neck dissection have been widely documented in the literature.9 -11
This study builds on prior work using artificial intelligence (AI) to make predictions about lymph node metastases in PTC from the primary tumor histopathology alone. Machine processes that mimic human cognition are broadly termed AI. One type of AI focuses on mimicking learning, by having computers perform tasks they were not explicitly programed to do and improving that performance with each iteration. This field of machine learning has many subsets, including deep learning, which uses multiple layered algorithms. Most analyses of visual data rely on a type of deep learning called convolutional neural networks (CNN). These algorithms perform analyses using interconnected processing nodes, hence the comparison to the nervous system, and have broad applications in clinical medicine. There have been numerous demonstrations in the literature of the ability of deep learning algorithms to make clinical predictions and diagnoses from all kinds of data, including visual radiology and histopathology.12 -22 The objective of this study was to further develop a CNN to predict lymph node metastases in PTC from multi-institution primary tumor histopathology data alone.
Methods
This study was performed in accordance with the Institutional Review Board requirements at the University of New Mexico (UNM) (Albuquerque, New Mexico) and the University of Kentucky (UK) (Lexington, Kentucky).
At both institutions, all cases of total or hemi-thyroidectomy with or without neck dissection between December 1, 2010 and December 31, 2019 were identified from pathology reports. Patients were excluded if they lacked complete pathology data or did not have majority classical or conventional papillary thyroid carcinoma (PTC) on final pathology. Only patients with pathologically confirmed lymph node status (positive or negative) were included. Cases were designated positive for lymph node metastases if there were at least 5 or more positive lymph nodes on the final pathology report. These criteria were based on risk factors for structural disease recurrence as outlined in the 2015 American Thyroid Association (ATA) Guidelines. 3 Negative cases were required to have at least 3 negative lymph nodes and no positive lymph nodes on the final pathology report in order to minimize false negatives given the retrospective study. All other patients were excluded. Based on protocols from the College of American Pathologists, a lymph node with any amount of tumor was defined as positive. 23 All identified cases were reviewed by a pathologist prior to final inclusion to verify the original pathology report. Ultimately, 418 cases of PTC were included in the study, 224 with regional metastases (positive) and 194 without (negative). Two hundred forty-four cases were from UK and 174 cases were from UNM.
The cases from both UNM and UK were scanned with the Aperio VERSA 200 slide scanner (Leica, Wetzlar, Germany) at 40x magnification. The cases from UNM were imported into a computer containing a 12 core, 2.2 GHz Intel Xeon Processor E5-2650 chip and a Nvidia Titan XP graphics card. The UK cases were imported into a computer containing a 3.6 GHz Intel Xeon Processor W-2123 chip and a Nvidia Quadro P400 graphics card. HALO-AI image analysis software (Indica Labs, Albuquerque, NM) was used to perform training and testing. HALO-AI uses a fully convolutional version of the VGG architecture with padding removed.
Representative tumor regions on a single slide with well-preserved histology were included for analysis in each case, after review and approval by the attending study pathologists. Annotations included the tumor and a transition zone of normal thyroid-tumor border including stroma. Three different algorithms were trained and tested from this data. First, the published UNM algorithm was tested on the cases from UK. 24 Second, an algorithm was developed from the UK cases and then tested on the cases from UNM. Lastly, a multi-institution algorithm was developed from cases from both institutions and then tested on the remaining cases in the combined data set.
The first algorithm (UNM algorithm on UK data), as previously reported, was trained on 95 randomly selected UNM cases (57 positive, 38 negative). 24 This was then tested on all 244 UK cases (123 positive, 121 negative). The second algorithm (UK algorithm on UNM data) was trained on all 244 UK cases and tested on all 174 UNM cases (101 positive, 73 negative). The third algorithm (combined data) was trained on 250 randomly selected cases from the combined UNM and UK data (130 positive, 120 negative) and tested on the remaining 168 cases (94 positive, 74 negative).
For all algorithms, analysis was performed by HALO-AI on PTC cases labeled as either positive or negative according to the criteria described above. The training cases were analyzed by HALO-AI as “image patches” of 400 × 400 pixels (where 1 pixel = 1 μm), at a resolution corresponding to a 5.5x or 2.5x digital view magnification.
Within the previously annotated tumor regions, the image patches were analyzed by HALO-AI were generated by automated selection of random points and cropping a patch around the point. These patches were further augmented with random rotations and random shifts to hue, saturation, contrast and brightness. Training was performed for a total of 70 193 analytic iterations for the first algorithm (UNM algorithm on UK data), 807 457 analytic iterations for the second algorithm (UK algorithm on UNM data), and 709 129 analytic iterations for the third algorithm (combined data). The training for all 3 algorithms used RMSProp (delta of 0.9), an optimization algorithm designed for neural networks, with a learning rate of 1e−3 reducing the learning rate by 10% every 2k iterations and an L2 regularization of 5e−4.19. During these iterations, the algorithm would change the node-weighted values continuously, based on the gold standard clinical lymph node status. The HALO-AI operator stopped the algorithm once an error rate/cross entropy rate of 0.1 or less was achieved, resulting in an error rate/cross entropy rate of 0.008 for the first algorithm (UNM on UK), 0.1 for the second algorithm (UK on UNM), and 0.070 for the third algorithm (combined data).
When testing the developed algorithm, HALO-AI analyzed the cases blindly, assigning each region (patch) a likelihood score for that area, which corresponded to the most probable assessment of nodal status. HALO-AI only analyzed the relevant annotated tumor regions, as was done during training. Because the analysis is based on patches, the output for each test case is X% area positive versus X% area negative, as demonstrated in Figure 1. These percentages are termed Area Distribution (AD) in this paper, with each test case receiving a positive AD (% of the case favored to represent positive lymphadenopathy) and a negative AD (% of the case favored to represent negative lymphadenopathy). Regions were labeled red if the algorithm predicted positive lymph node status, while regions were labeled green if the algorithm predicted negative lymph node status. Thus, the result for each case is a mosaic of positive and negative regions according to the AI algorithm.

Representative tumor slides from the University of New Mexico (UNM) (A) and the University of Kentucky (UK) (B) are shown. The slides are annotated with a green line marking the tumor and a berth of normal tissue around it. When testing the algorithm, the convolutional neural network (CNN) highlights the positive Area Distribution (AD) in red as shown on these representative histopathology slides. The rest of the slide is marked green by the algorithm. The positive AD was used in statistical analysis to evaluate the algorithm’s success and determine the ideal thresholds.
Population differences between UNM and UK data sets, as well as between the positive and negative cohorts, were analyzed with simple t-tests as most comparable measurements were proportions. Continuous variables like age were also analyzed with t-tests given the large data set and normal distribution. Receiver operating characteristic (ROC) curves were used to analyze the diagnostic capabilities of various positive AD percentage thresholds for each algorithm. This method plots the sensitivity against 1-specificity for a range of possible thresholds in order to identify the most useful diagnostic threshold in continuous clinical data. A Youden J statistic was used to find the ideal threshold that maximized sensitivity and specificity and this optimal threshold was used to calculate a positive predictive value (PPV) and negative predictive value (NPV) for each algorithm. The area under the curve (AUC) was calculated for each ROC curve, which is a simple measure of the predictive ability of the model. The closer the AUC is to 1, the better the capability of the model to distinguish between the categories in question.
Results
Descriptive statistics for the study population, compared by institution and nodal status, can be found in Table 1. The UNM population was slightly older (49.2 vs 44.9 years, P < .01) than the UK population. The UNM population was also significantly less white (59.2% non-white vs 4.5% non-white, P < .01). As expected, the population with positive lymph node metastases was different from the population without lymph node metastases. The group with lymph node metastases had fewer female patients (64.7% vs 77.3%, P < .01), was younger (44.8 vs 48.9 years, P < .01), and more likely to be non-white (33.5% vs 20.1%, P < .01). They also had larger primary tumors (30.6 vs 18.4 mm, P < .01) and a higher likelihood of both positive margins (32.1% vs 9.8%, P < .01) and aggressive pathology features, such as lymphovascular invasion or extrathyroidal extension (82.6% vs 27.3%, P < .01).
Descriptive Population Statistics, by Nodal Status and by Institution.
The diagnostic capabilities of the first algorithm (UNM algorithm on UK data) were maximized at a positive Area Distribution (AD) percentage threshold of 68%, for which the maximum Youden J statistic was 0.21. This means that our model was most successful at predicting nodal metastases when a cut-off of 68% positive AD was used. At these thresholds, if the algorithm called greater than 68% of the slide positive, the case was deemed positive. Likewise, if the algorithm called less than 68% of the slide positive, the case was deemed negative. At this positive AD percentage threshold, the algorithm had a sensitivity of 74% and a specificity of 47%. In this study, the positive predictive value (PPV) and negative predictive value (NPV) were 59% and 64% respectively for the first algorithm. The full receiver operating characteristic (ROC) curve for this method is shown in Figure 2A. The area under the curve (AUC) was 0.59 for the first algorithm.

Receiver operating characteristic (ROC) curves for the first algorithm (UNM algorithm on UK data) (A), the second algorithm (UK algorithm on UNM data) (B), and the third algorithm (combined data) (C). The ideal thresholds for each algorithm are marked with a red dot. The ideal positive area of distribution (AD) threshold, sensitivity, specificity, and area under the curve (AUC) for each algorithm is shown.
The diagnostic capabilities of the second algorithm (UK algorithm on UNM data) were maximized at a positive Area Distribution (AD) percentage threshold of 42%, for which the maximum Youden J statistic was 0.26. At this positive AD percentage threshold, the algorithm had a sensitivity of 65% and a specificity of 61%. In this study, the positive predictive value (PPV) and negative predictive value (NPV) were 74% and 59% respectively for the first algorithm. The full receiver operating characteristic (ROC) curve for this method is shown in Figure 2B. The area under the curve (AUC) was 0.64 for the second algorithm.
The diagnostic capabilities of the third algorithm (combined data) were maximized at a positive Area Distribution (AD) percentage threshold of 60%, for which the maximum Youden J statistic was 0.59. At this positive AD percentage threshold, the algorithm had a sensitivity of 68% and a specificity of 91%. In this study, the positive predictive value (PPV) and negative predictive value (NPV) were 90% and 69% respectively for the first algorithm. The full receiver operating characteristic (ROC) curve for this method is shown in Figure 2C. The area under the curve (AUC) was 0.84 for the third algorithm.
Discussion
Artificial intelligence (AI) has the potential to revolutionize clinical medicine with its ability to generate novel diagnostic information that is beyond current human capabilities. Clinical pathologists do not predict metastases from review of primary tumors alone, but the CNN in this multi-institutional study completed this task with good success. This multi-institution, proof of concept study is the next step of many in bringing this type of AI technology to clinical use.
The most accurate AI algorithm was trained on combined data from both UNM and UK, with a clinically viable area under the curve (AUC) of 0.84. Looking forward to clinical applications, it would be important to demonstrate that an algorithm generated in one setting could still be effective when applied to a different setting. Unfortunately, neither single institution algorithm performed that well, likely due to their smaller data sets which were presumably insufficient to overcome institutional variation in the visual data, such as differences in staining or storage techniques between institutions. However, even with the smallest data set, the algorithm did better than a coin flip, with an AUC of 0.59, when tested on completely distinct data from a different institution. Ideally, a true test of broad applicability would involve an algorithm trained on heterogenous multi-institution data and then tested on a separate heterogeneous data set from different institutions, which require more case numbers. The success of all 3 algorithms improves as the size of the training data set increase. If the presence of lymphovascular invasion or extrathyroidal extension were used to predict lymph node metastases in this data set, the specificity would be only 78.3% compared to a specificity of 91% in our best performing algorithm. This specificity is particularly useful for this clinical question, where the algorithm’s utility is in upstaging intermediate risk patients to radioactive iodine ablation.
In addition to the number of data points, the number of analytic iterations is crucial to assessing whether a deep learning algorithm has overfit or underfit the data. The CNN software used in this study has 2 inherent processes designed to optimize this. First, it applies a separate optimization algorithm to ensure a responsible and deescalating learning rate as the iterations of the algorithm increase. 25 Second, as detailed in the methods, it stops the algorithm once the error/cross entropy rate is below 0.1. It does not try to overfit the data by minimizing this value.
Supervision refers to the amount of human data curation that is done prior to algorithm training. In the prior example, the highly supervised training involved a pathologist circling just the tumor on the slide, imposing their preconceptions about which parts of the slide would be important. In the more weakly supervised training that was used in this study, a much more generous portion of the slide was shown to the algorithm and it performed better. Weakly supervised algorithms have shown success in making similar predictions, such as predicting regional mutational status from histopathology slides. 21 The annotation used in this study entailed moderate supervision, as it included a cuff of normal tissue around the tumor. This method, as opposed to whole slides or a section of tumor alone, performed best in our prior work. 24 The inclusion of the tumor transition zone in the best performing models, seems to suggest some sort of biologic process at the tumor edge that is invisible to humans. Hypotheses like this may be useful in informing future research.
While the emergence of successful machine learning algorithms for clinical medicine is exciting, some hesitancy about their widespread application is reasonable. Biases in data and analysis are valid concerns. 26 Kleppe et al outline a multi-stage process to responsibly bring machine learning algorithms to clinical fruition. This study exists between phases 1 and 2 out of the 6 phases on their timeline, firmly in development and testing, often with redevelopment and retesting. The best combined data algorithm in this study is still in need of additional external validation of its performance, with larger data sets from additional institutions as well as the more rigorous bootstrapping or cross-validation that is possible with larger data sets.
Another broad concern involves the true clinical utility of some proposed algorithms, including this one. How much benefit could this additional prognostic algorithm truly add to the care of a patient with papillary thyroid cancer? Our data set is retrospective and suffers from selection bias given the inclusion requirement of proven negative lymph nodes on final pathology, which is complicated by low rates of routine central neck dissection during thyroidectomy. Since this was a retrospective study, there was no standardization of lymph node sampling during thyroidectomy. Stringent criteria were used for negative cases to ensure 2 distinct groups for analysis. In addition, even though we only included histopathology from the primary tumor in our analysis, the clinical decision-making context for the patients in this study was variable. Future prospective studies could include only patients with N0 necks who all receive selective central neck dissection.
Our own data also show the myriad clinical and pathologic predictors for regional metastases, as evidenced by the statistical and clinical differences between the positive and negative cohorts. While none of these risk factors are truly independent predictors of lymph node metastases, they do reduce the marginal utility of targeted predictive algorithms like the one in this study. The best algorithms will incorporate these predictive factors, combining visual, pathologic, radiologic, and clinical data. This study helps validate the utility of visual histopathology data as an effective predictor of clinical outcomes and is a necessary step in building these powerful, complex algorithms. The final 2 stages of the Kleppe timeline focus on demonstrating the clinical utility, not just performance ability, of machine learning algorithms. These demonstrations require prospective data with specific expected clinical outcomes. These prospective cohorts also open up the possibility of frozen section or cytopathology data to be included, in addition to improving the algorithm with a homogenous study population all undergoing the same treatment. The future of AI algorithms in medicine is likely personalized medicine, where patients and their surgeons can see reliable predictions of clinical or surgical outcomes based on all of the available data.
Conclusion
A convolutional neural network can produce an accurate and robust algorithm that is capable of predicting nodal metastases from primary PTC tumors alone. The best performing algorithm in this study had a sensitivity of 68% and a specificity of 91%. Larger prospective data sets from additional institutions will be required to perfect the algorithm and improve patient care.
Footnotes
Acknowledgements
We would like to thank Fred Schultz, an analyst in the Department of Pathology and the Human Tissue Repository at the University of New Mexico, for his technical work on this project.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
