Abstract
A subset of patients with neuroblastoma are at extremely high risk for treatment failure, though they are not identifiable at diagnosis and therefore have the highest mortality with conventional treatment approaches. Despite tremendous understanding of clinical and biological features that correlate with prognosis, neuroblastoma at ultra-high risk for treatment failure remains a diagnostic challenge. As a first step towards improving prognostic risk stratification within the high-risk group of patients, we determined the feasibility of using computerized image analysis and proteomic profiling on single slides from diagnostic tissue specimens. After expert pathologist review of tumor sections to ensure quality and representative material input, we evaluated multiple regions of single slides as well as multiple sections from different patients’ tumors using computational histologic analysis and semiquantitative proteomic profiling. We found that both approaches determined that intertumor heterogeneity was greater than intratumor heterogeneity. Unbiased clustering of samples was greatest within a tumor, suggesting a single section can be representative of the tumor as a whole. There is expected heterogeneity between tumor samples from different individuals with a high degree of similarity among specimens derived from the same patient. Both techniques are novel to supplement pathologist review of neuroblastoma for refined risk stratification, particularly since we demonstrate these results using only a single slide derived from what is usually a scarce tissue resource. Due to limitations of traditional approaches for upfront stratification, integration of new modalities with data derived from one section of tumor hold promise as tools to improve outcomes.
Neuroblastoma, a neural crest cell derived cancer, is the most common extracranial solid tumor of childhood. 1 Aggressive disease, referred to as high-risk neuroblastoma, is treated with intensive multimodal therapy, yet responses are suboptimal. Patients with neuroblastoma at ultra-high risk for treatment failure, a subset with the worst prognosis, encompass up to 15% of those with high-risk disease and is defined as death from disease within 18 months of diagnosis despite conventional treatment. 2 Extraordinary clinical, molecular, and histologic heterogeneity of neuroblastoma have made it challenging to identify prognostic biomarkers at diagnosis for those with the poorest outcomes. Early identification of this risk group of patients would provide an opportunity to offer alternative therapies that have potential to improve outcome.
Clinically, neuroblastoma is remarkable in that the disease can be highly aggressive and metastatic at diagnosis requiring intensive multimodal therapy to achieve only modest outcomes, or it can present with substantial tumor burden that spontaneously regresses. These neural crest cell derived tumors arise most frequently from the adrenal medulla, but can occur anywhere along the sympathetic chain. While the median age at diagnosis is 18 months, neuroblastoma can be diagnosed prenatally and rarely in adolescent and young adult patients.3,4 In general, children less than 18 months have superior outcomes. Molecular and genetic alterations are also variable. MYCN amplification, diploid DNA content, and specific segmental chromosomal aberrations predict poor outcome.5–9 In addition, more recently, ALK aberrations have been correlated with inferior outcomes in 14% of high-risk patients. 10 Understanding of histologic intertumor heterogeneity has also revealed features that have negative prognostic value, including increased mitosis-karyorrhexis index and lack of neuroblast differentiation, and features that have favorable prognostic value such as Schwannian stroma dominance.11,12 Emerging histologic markers such as presence of prominent nucleoli may also become integrated into risk classification in future studies. 13 Many of these clinic-pathologic prognostic variables form the basis of the current risk classification used in neuroblastoma 5 ; yet, limitations exist as more than half of patients with high-risk disease do not survive despite highly intensive multimodal therapy. Therefore, identification of the patients who will recur early are refractory to upfront treatment and/or die from disease within 18 months (i.e., neuroblastoma at ultra-high risk for treatment failure) requires exploration of novel approaches using limited diagnostic tumor material. As a first step, we set out to confirm that sectional analyses are reflective of whole tumor based on 2 technical assays.
Computer-aided diagnosis is increasingly being studied in malignancies, including medulloblastoma 14 and prostate cancer. 15 We have reported on the applicability of this technology in neuroblastoma16–18; yet, it remains underutilized and has not been applied to improve diagnostic subclassification of high-risk disease. Such techniques can extract features with high throughput and high sensitivity and therefore supplement human interpretation of histopathology. Further, spectroscopic-based proteomic analysis of scarce formalin-fixed paraffin-embedded tissue is increasingly being applied in the study of cancer.19–21 Proteomic profiling has identified clinically relevant proteins, including high-risk neuroblastoma,22–24 though not to differentiate the most highly aggressive disease. In a series of patient tumors, we used automated image analysis tools and ion intensity-based label-free semiquantitative proteomics to evaluate histoproteomic data from multiple sections of a tumor and multiple regions of a slide. We then compared individual samples between and within tumors and found that intratumor heterogeneity is present; however, it is less than the degree of intertumor heterogeneity. This supports the use of single section analysis, when chosen with the guidance of an expert pathologist, as broadly representative of the entire tumor for computational histologic and proteomic analyses. These data can then be used to support the development of new variables that could have prognostic and potentially therapeutic significance for patients with neuroblastoma at ultra-high risk for treatment failure.
Materials and Methods
De-identified formalin-fixed paraffin-embedded unstained tumor material on glass slides from patients with high-risk neuroblastoma were obtained from the Children’s Oncology Group clinically annotated Biospecimen Bank. Institutional review board exempt status was granted. For this study, we only included high-risk neuroblastoma patients defined as greater than 18 months and International Neuroblastoma Staging System stage 4 (presence of distant metastases). Samples were from patients who had either neuroblastoma at ultra-high risk for treatment failure, which was defined here as death from disease in less than 18 months, or high-risk disease that was successfully treated (survival without recurrence for greater than 5 years). For each slide studied, we had at least 3 sequentially cut slides so that hematoxylin and eosin (H&E) staining would overlap with region of tumor for proteomic analysis. A board-certified pediatric pathologist with expertise in neuroblastoma (bootstrap probability) confirmed histologic diagnoses and highlighted the best preserved and most representative regions of tumor sections (eg, nodules in ganglioneuroblastoma) for subsequent analyses. As represented in Figure 1, multiple nonadjacent sections from different blocks of 5 patients with known outcome were used for analysis (Table 1). Slides in Set E were of determined by expert pathology review to be of insufficient quality for automated image analysis.
Representative sectioning of a patient tumor sample, eg, Patient Set C. Tumors or tumor biopsies are sectioned into blocks and slides are made from sequential cuts of formalin-fixed paraffin-embedded tissue. The numbers correlate with “Tissue Section ID” when applicable, the dotted line reflects the section that was used for the slide, and the histologic images correlate with multiple ROIs per image. An immediately adjacent slide was used for proteomic analyses. Correlative Clinicopathologic Covariates for Patients With Multiple Sections per Tumor. Listed are known variables currently used for risk stratification of neuroblastoma. INPC takes into account tumor histology with degree of differentiation. INSS, international neuroblastoma staging system; MKI, mitosis-karyorrhexis index; INPC, international neuroblastoma pathology classification; NA, non-amplified; A, amplified; DI, DNA index; GNB, nodular ganglioneuroblastoma; -, data not available. Each Tissue Section ID reflects a distinct nonadjacent section of tumor (see Figure 1).
Automated Image Analysis: Distribution of 88 ROI Images Into 5 Different Patient Sets, With Correctly Matched Sections.
Thirty-one tissue sections were extracted from 26 neuroblastoma patients. Sets A–D are composed of multiple nonadjacent tumor sections for 4 individual patients. Set F includes at least 1 ROI from 22 distinct patients. Correctly matched sections were based on 1 Patient Set compared internally as well as to all others combined. ROI, region of interest.
Automated Image Analysis
As a first step in image analysis, we performed color deconvolution
25
to separate hematoxylin, which primarily stains nuclei, from eosin in each ROI image. The texture of the extracted nuclei was further analyzed by 2 filter banks.26,27 One is rotationally symmetric which makes it suitable for encoding circular texture patterns,
27
while the other is effective at encoding texture comprising of linear structures.
26
We further computed the morphological features resulting from the segmentation of the filtered images. To encode the geometrical arrangement, we computed the average inter-centroid distances between the nuclei. This is followed by message passing clustering
28
to find the optimal number of clusters in each ROI image. We then used the centroids of these clusters as a proxy for the ROI image. Each centroid is a feature vector which encodes certain aspects of the ROI image. For evaluation, we first convert the test image into feature vectors. Then, we compare the closeness of these feature vectors with the centroids of each ROI image. The ROI image that results in a minimum mean square error is considered as a closest match for the test image. Therefore, the proposed image analysis method is self-contained; once it is trained, it requires input images and all the algorithm steps are executed on these images. Further, the clustering is carried out to convert each ROI image into a set of centroids rather than to cluster images. As these centroids lie in a very high dimensional feature space, it is practically impossible to graphically show the distribution of data points. Figure 2 shows the block diagram of this process.
Block diagram of the image analysis workflow. The circular filter bank extracts circular features while the linear filter bank encodes linear texture features. C1, C2, C3, C4, and Cn represent the cluster centroids resulting from a message passing clustering algorithm.
Proteomics
Using 5- to 10 -µm-thick formalin-fixed paraffin-embedded tumor sections on a glass slide, antigen retrieval methods were used to extract one microgram of protein from the tissue. 29 Proteins were digested with trypsin using Filter Aided Sample Preparation as previously described and peptides were desalted using StageTips prior to LC-MS/MS analysis.30,31 Proteolytic peptides were separated on a 60 cm × 75 µm column packed in house with 4 µm C12 Jupiter Proteo beads (Phenomenex, CA, USA) that were maintained at 50℃. Peptides were separated at a flow rate of 300 nL/min with a 90-min gradient of 2% to 42% acetonitrile with 0.5% acetic acid. Eluted peptides were detected with an LTQ Orbitrap Hybrid Mass Spectrometer (Thermo Scientific, MA, USA) with a full mass range of m/z = 370–2000. The 5 most intense precursor masses were selected for fragmentation in the linear ion trap.
Proteomics data were searched against the human UniProt database with MaxQuant v1.3.0.5 using a 1% false discovery rate. Identified proteins with nonsignificant mapping (P value > .05) and high missing rate across samples (missing data ≥ 0.2) were excluded from analysis. Data were then normalized using central tendency measure and missing data were imputed using singular value decomposition using R functions from InfernoRDN software (v1.1.5438, http://omics.pnl.gov/software/infernordn).32,33 Two samples (from individual sections) were excluded due to high data missing rate and low correlation with other samples. Remaining samples were clustered using the pvclust (v1.3-2) R package. 34 We used hierarchical clustering with an average agglomeration method and correlation distance measure. To determine stability and significance of the clustering results, we performed bootstrap resampling of the data with 1000 replications.
Results
To explore the degree of histoproteomic heterogeneity within and between primary tumor specimens, we applied 2 complementary methods that are currently not integrated in neuroblastoma classification. Results from computerized microscopy and proteomic profiling were obtained from analysis of multiple sections of tumors (Table 1 and Figure 1).
Automated Image Analysis
We defined histologic heterogeneity in terms of nuclear texture, nuclear morphology, and geometric arrangement of the nuclei. These features are designed to characterize underlying tissue characteristics: eg, geometrical arrangement is a measure of nuclear crowding while nuclear texture and morphology distinguish different cell types. The results are presented in terms of correctly matched tissue sections, ie, if an ROI image from a tissue section is correctly matched to an ROI image of a different tissue section from the same patient. A different strategy was used for Set F as it contained single tissue sections from each patient: we considered an ROI image to be a correct match if the closest ROI image is from the same tissue section. The results are presented in Table 2 (right column) and show that tissue sections in Set A, B, C, and D were all correctly matched to each other while only 3 tissue sections in Set F correctly matched to other tissues, indirect confirmation of intertumor histologic heterogeneity.
Proteomics
The 1495 most abundant proteins were evaluated across samples. Proteins with a high rate of missing data (≥ 0.2) across all samples were excluded, leaving 839 proteins. Proteome wide peptide detection allowed for unbiased hierarchical clustering analysis on all samples. We used average clustering criteria, correlation methods for distance, and pairwise covariance measures. Confidence of clusters was measured using approximately unbiased and bootstrap probability (1000 replications) methods to identify clusters with > 95% probability. Clustering dendrogram revealed that intratumor distance was significantly shorter than the intertumor distance (Figure 2), indicating samples from the same tumor were more closely related than samples between tumors. When testing only subjects with multiple tumor samples, 4 out of 5 samples from the same subject clustered with > 95% confidence (Figure 3). The samples from the remaining subject have a confidence >80% but differing rates of missing data. When using all samples available, tumor samples from the same subject still clustered together and had a greater correlation than tumor samples from different subjects (Figure S1). A validation cohort of 7 distinct sections from separate blocks of tumors of 3 high-risk patients confirmed that tumors cluster by protein expression (data not shown).
Cluster dendrogram reveals a high degree of intratumor similarity as assessed by proteomic profiling. Clustering dendrogram of 11 distinct sections from 5 tumors with AU (approximately unbiased, red) and BP (bootstrap probability, green) values (%). A higher AU and BP value indicates significant clustering. Sample number corresponds to Tissue Section ID, and consecutive number pairs are different sections from the same tumor. Red box ≥95% confidence clustering.
Discussion
Decades of coordinated international effort to contribute to tumor biobanks have enabled development of scientifically valuable neuroblastoma tissue biorepositories with samples available for researchers. However, even though neuroblastoma is among the more common childhood cancers, the rarity of certain subsets, including those with neuroblastoma at ultra-high risk for treatment failure, means it can take years to acquire even dozens of cases. Further, small biopsies and finite precious study material mandates limits to exploratory research on such specimens. We set out to apply novel techniques to small samples of tissue to better understand neuroblastoma biology, focusing on automated image analysis and proteomic profiling of high-risk neuroblastoma. Specifically, our ultimate goal is to improve prognostic risk stratification at the time of diagnosis and therefore predict patients who will have the poorest outcomes. As no molecular, histologic, or clinical feature has accomplished this to date, our new approaches have the potential to alter the paradigm of neuroblastoma risk classification by leveraging machine learning and –omics data from a single slide of fixed tumor material. We demonstrate that pathologist-determined representative tumor sections are reliable surrogates for whole tumor analysis for the described technical assays. These data should inform additional –omics strategies from preserved tumor material, particularly when such material is in short supply.
The International Neuroblastoma Pathology Classification developed by Shimada et al. 11 is a histologic tool for diagnostic risk stratification that takes into account tumor differentiation and mitosis-karyorrhexis index. It helps to categorize patients into very low-, low-, intermediate- or high-risk group based together with other clinicobiologic features, and risk stratification is used to determine treatment intensity. There are 2 main limitations to the current histology methods: First, it is a tedious grading scale with significant inter- and intra-observer variability; second, it does not include features that distinguish patients with neuroblastoma at ultra-high risk for treatment failure. Automated image analysis and digital pathology are increasingly being studied across disciplines and diagnoses. 35 Integration of such technology into current diagnostic workups has the potential to augment traditional histologic findings as it can consistently and reliably extract histologic features of a tumor to supplement expert pathologist review for improved diagnostic classification of disease. The aims of image analysis are (1) to provide a mathematical model to define intratumor heterogeneity, which is difficult for pathologists to quantify, and (2) to provide a modality to compare the findings with those from proteomics in a quantitative manner. The difference of this mathematical model from a human pathologist is that while the pathologist would be able to qualitatively differentiate heterogeneity, it may be difficult to describe in words the features used to assess this. Additionally, different pathologists may have varied interpretation of heterogeneity. The current image analysis method is an effort towards minimizing quantification and consistency challenges.
Our results demonstrate that the nuclear texture, morphology, and geometrical arrangement remains relatively similar across different tissue sections of the same tumor from one patient. Using complementary samples, a global proteomic analysis also reveals profiles that are more uniform within tumors as compared to between tumors. Proteomics is known to be a less robust molecular analysis tool, especially as compared to genomics approaches, as the replicate variability is fairly high. Here, we did not aim to achieve a definitive proteomic profile but rather establish degree of heterogeneity based on such profiling. Therefore, while intratumor heterogeneity exists in different sections of the same tumor based on the described techniques, it is less than intertumor variability. This anticipated finding is consistent with reported limited intratumor heterogeneity of the MYCN oncogene.36,37 Moreover, the results of Set F also depict minimal intra-section heterogeneity, as the ROI images from the same section were correctly matched to each other. From this perspective, we provide rationale to use single tissue section proteomics and image analysis in ongoing research to define biologic drivers of the most highly aggressive neuroblastoma, a clinically important subgroup of patients who have highly lethal disease, and for whom no diagnostic biomarkers exist. Having confidence that tumor material from a section is reflective of whole tumor is essential to proceed with additional biomolecular analyses of tumor.
Our work also further establishes the rationale for development of a multiresolution and multiclassifier approach to deal with the large size and complexity of digital images and to mimic pathologist review.38–40 We can accomplish this by using an experienced pathologist who determines the best preserved and most representative section of primary tumor as a surrogate for the whole specimen when tissue is scarce. Integration of automated image analysis and proteomics approaches, together with other well-studied features such as age and MYCN amplification, may establish a prognostic classifier that defines neuroblastoma at ultra-high risk for treatment failure as a distinct subset of high-risk patients, and this methodology could then be applied to individual patients who present with high-risk neuroblastoma.
Our study is the first to use automated microscopy to evaluate tumor histologic heterogeneity. This is an important first step towards integration of such techniques into contemporary diagnostic clinicopathologic risk classification in neuroblastoma which currently cannot stratify patients at ultra-high risk for treatment failure. A limitation of the current study is that the findings are derived from a focused sample set—a relatively small number of high-risk cases—and future work will involve larger sample sets of high-risk disease where there is greatest unmet need to improve outcomes. Further, analysis of single tumor sections will invariably miss some degree of tumor heterogeneity that may have biologic relevance, such as nodular components in ganglioneuroblastoma and in the case of metastases in a single patient.41,42 Additional annotation of future samples will include histologic diagnosis of neuroblastoma versus ganglioneuroblastoma in order to assess greater expected heterogeneity between these samples rather than within. Also, formalin-fixed paraffin-embedded samples are variably processed at institutions, which may contribute to the intertumor differences we identified, though the intratumor differences are substantially less. After additional retrospective analyses, we intend to apply digital microscopy and proteomics as exploratory or integrated prognostic biomarkers in clinical studies in order to refine risk stratification to include an additional higher risk subset of neuroblastoma. Importantly, we will further examine the proteomics data to gain biologic insight about the pathogenesis of this highly aggressive disease. Improved diagnostic accuracy will result in optimized treatment approaches, particularly for patients who have an otherwise dismal prognosis with conventional therapy.
Footnotes
Acknowledgments
We thank the Children’s Oncology Group for providing patient samples, and Dr. Michael Hogarty for his guidance and insight.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Funding was provided by KL2 RR024132 (DW), National Cancer Center (DW), Bear Necessities/Rally Foundation (DW), U24CA199374 (MG), U24CA196173/114766 (COG Biopathology Center), and U10CA180899 (COG Statistics and Data Center). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Cancer Institute or the National Institutes of Health.
