Abstract
In Clinical Epidemiology, receiver operating characteristic (ROC) analysis is a standard approach for the evaluation of the performance of diagnostic tests for binary classification based on a tumour marker distribution. The area under a ROC curve is a popular indicator of test accuracy, but its use has been questioned when the curve is asymmetric. This situation often happens when the marker concentrations overlap in the two groups under study in the range of low specificity, corresponding to a subset of values useless for classification purposes (non-informative values). The partial area under the curve at a high specificity threshold has been proposed as an alternative, but a method to identify an optimal cut-off that separates informative from non-informative values is not yet available. In this study, a new statistical approach is proposed to perform this task. Furthermore, a statistical test associated with the area under a ROC curve corresponding to informative values only (restricted ROC curve) is provided and its properties are explored by extensive simulations. Finally, the proposed method is applied to a real data set containing peripheral blood levels of six tumour markers proposed for the diagnosis of neuroblastoma. A new approach to combine couples of markers for classification purposes is also illustrated.
Keywords
1 Introduction
Receiver operating characteristic (ROC) curves are statistical tools largely applied in Clinical Epidemiology to evaluate the performance of tumour markers (TMs) in the context of binary classification problems.1–3 Because the early steps for a marker validation, in general, require a case-control approach,2,4 in this article the class of diseased subjects will be referred to as the ‘cases’ and the referent group, which typically includes healthy individuals or subjects affected by less severe diseases, as the ‘controls’. In general, to be useful for diagnostic purposes a TM should present, on average, higher values in the class of the cases than in that of the controls. A binary test may be obtained by selecting a specific value of a continuous TM and by defining as positive any test with value exceeding such a threshold, and as negative the remaining other tests. The proportion of positive cases provides an estimate of the test sensitivity and the proportion of negative controls an estimate of its specificity. A ROC curve is obtained by plotting the true positive fraction (sensitivity) vs the false positive fraction (1-specificity) using all the available values of a TM concentration. Figure 1(a) shows a theoretical ROC curve obtained from infinite values of a hypothetical TM and an empirical ROC curve obtained from a finite sample of 200 units (100 cases and 100 controls) randomly selected from the same binormal distribution (displayed in panel (b)).
Theoretical and empirical ROC curves (panel (a)) and a corresponding density probability distribution (panel (b)) (binormal model with the same variance). The empirical curves was obtained from 200 random samples (100 cases and 100 controls). FPF: false positive fraction; TPF: true positive fraction.
The area under a ROC curve (AUC) is a popular measure of diagnostic accuracy. It is equivalent to the Mann-Whitney U statistic, thus representing the probability that a subject, randomly selected among the case class, shows a marker value higher than a subject randomly extracted from the controls.5,6 For completely non-informative TMs the ROC curve will approach the rising diagonal (called ‘chance diagonal’ or ‘chance line’, Figure 1(a)) and AUC will tend to 0.5, i.e. the expected probability for a classification due to chance alone. On the contrary, in the case of a perfect classification the ROC curve will reach the point of the highest theoretical accuracy (sensitivity and specificity both 100%) and AUC will tend to one, i.e. the highest probability value. A statistical test to verify that AUC is different from its expected value under the null hypothesis is easily performed exploiting the normal asymptotic distribution of the U statistic7–9
An optimal cut-off for a binary classification test may be identified on a ROC curve as the value corresponding to the highest vertical distance J from the chance line (Figure 1, panel (a)). J also corresponds to the highest value of the Youden’s index, a popular measure of pure accuracy.2,11,12
Proper ROC curves are concave and symmetric and never cross each other, thus allowing an easy and reliable comparison between the corresponding markers. In fact, for proper ROC curves, the highest AUC ever corresponds to the highest sensitivity at any specificity value. Furthermore, the selection of an optimal cut-off on proper curves is easily performed, because they include a unique point that corresponds to the optimal threshold based on the Youden’s index.2,11,12 However, in the analysis of TM, asymmetric (not proper) ROC curves are often encountered, due to the presence in the two classes under study of overlapping marker values at an extreme of their distributions. In many situations, a lack of sensitivity is observed in the correspondence of the lowest range of specificity values,
13
causing the corresponding ROC curve to be unimodal and right asymmetric with respect to the chance line. In this article, curves of this kind will be called ‘positive skewed ROC’. Four examples of such curves are illustrated in Figure 2(a) to (d), while Figure 3(a) to (d) shows four corresponding density distributions of four hypothetical TM concentrations. In more details, the ROC curve in Figure 2(a) is positive skewed and concave, and may correspond to a unimodal distribution of marker values among the controls and a bimodal ones among the cases (‘Normal–Binormal model’, Figure 3(a)). In the real world, this situation may occur when a subgroup of cases does not differ from the controls for the expression of the TM under study.
13
Another situation is that illustrated in Figure 3(b), corresponding to the ROC curve in Figure 2(b), in which cases and controls have a similar distribution density (both normal), but among the cases the marker have both higher mean and variance values (‘Binormal model’ with different variances). This behaviour may be considered as a variant of the previous one, when the two subgroups of cases are less distinguishable. In that case, the corresponding ROC tends to lose its concavity at the lowest scale of specificity and it may even cross the chance line2,3,13,14 (Figure 2(b)). Moreover, Figure 3(c) shows a bimodal distribution both in the cases and in the controls, with a subgroup within any class showing a similar TM distribution in the range of low marker values (‘Binormal–Binormal model’). This behaviour may be attributable to a detection threshold of the device used to measure the TM concentrations, which automatically replaces null values with a white noise. The corresponding ROC curve will be similar to that in Figure 2(c), when the curve approaches the chance line at low specificity values. Finally, a quite common variant of this situation is that of a TM with mass at zero, illustrated in Figure 3(d), in which the range of non-informative values is made by some zero values (‘Zero-inflated Binormal’ model). The corresponding ROC curve will be still positive skewed and it will also show a jump discontinuity of the derivative,
15
like the curve in Figure 2(d).
Example of four positive skewed ROC curves. FPF: false positive fraction; TPF: true positive fraction. Example of four density distributions corresponding to the positive skewed ROC curves in Figure 2. (a) Normal–binormal model; (b) binormal model with different variances; (c) binormal–binormal model; and (d) zero-inflated binormal model.

In the case of not proper ROC curves, some Authors have proposed the use of the partial area under the curve (pAUC) at an a priori selected cut-off of high specificity, or based on some utility function.13,16,17 All these approaches, however, do not provide a method to identify the range of non-informative values, if any, within the TM distribution. In general, subjects whose values fall in such a range are defined as ‘test negative’, albeit this range may include different proportions of cases and controls.
In this article, a new simple method of ROC analysis is illustrated, which allows to identify an optimal threshold to separate the range of non-informative values of a TM from those potentially useful for classification purposes. A statistical test associated with the ROC curve corresponding to the informative range only (restricted ROC, rROC) is also provided. Finally, a method to combine restricted and standard ROC analysis is illustrated and applied to a real data set of six TMs for the diagnosis of localized neuroblastoma.
This article is organized as follows: in Section 2 a new statistical test for rROC is provided, based on the standard normal distribution; in Section 3 its statistical power is investigated by extensive simulations and compared with that of the standard test on AUC; in Section 4 a simple rule for the choice between standard and rROC curves is illustrated; in Section 5 a method to combine information from standard and rROC curves is analysed; in Section 6 the new proposed method is applied to a real data set of concentrations of six putative TMs proposed for the diagnosis of neuroblastoma; in Section 7 a discussion of the method and the obtained results is provided.
2 Definition of a new statistical test for rROC
Let X be a TM whose concentrations are expressed on a continuous scale in two classes (controls and cases). Let Y be a random sample of X-values sorted in an ascending order in an array {yi} (i = 1, … , N) and let n0 be the sample size of the controls and n1 that of the cases, with n0 + n1 = N. A left rROC curve (simply called rROC, because right restriction will not be considered in this article) is defined as the ROC curve obtained after the elimination from Y of the first j ordered values. For any value of j ( j = 0, … , N), a rROC
j
may be identified and the corresponding area under the curve (rAUCj) estimated. Similarly to equation (1) a test statistic may be associated to rAUCj as follows
The new statistic test proposed in this article is defined as follows
Accordingly to equations (3) and (5), rzAUC represents the highest ‘standardized’ AUC, i.e. the value that identifies the rROC with the highest accuracy, estimated taking into account both the probability of a correct classification (i.e. rAUC) and a measure of its variability. Accordingly, the corresponding j allows the identification of a cut-off Cj (if any) which separates informative TM values from non-informative ones. j may be zero in the case of non-restriction, i.e. when the standard AUC has a higher performance than any other rAUCj. Finally, to allow rzAUC to be unique for any ROC curve, in the case of more than one value from equation (5), only that corresponding to the lowest j is retained, corresponding to the rROC based on the largest sample size.
The probability density of the rzAUC statistic under the null hypothesis H0 of an equal distribution of TM values among the two classes was investigated by extensive random simulations. N samples from a standard normal distribution were generated by letting N vary from 10 to 500 in each class. Each simulation was repeated 105 times in order to obtain stable estimates of the expected value and of the variance of rzAUC. The analysis was performed by a software ad hoc developed by the freely available package Microsoft Visual Basic Express 2010. Random values from an uniform probability density were obtained by the RAN1 algorithm, 18 while normal distributions were obtained by the GASDEV algorithm. 18 Furthermore, most routines for both standard and rROC analysis were implemented in an open source statistical program (called ‘rROC’) freely available at: http://www.ge.ieiit.cnr.it/∼muselli/supplement/smmr-2012.html. Finally, histograms and the quantile–quantile inverse normal plots (qqplots), for the evaluation of the normality of rzAUC distribution, were obtained by STATA for Windows statistical package (release 11.0, Stata Corporation, College Station, TX, USA).
Figure 1S(a) to (f) in Supplemental Material shows six histograms of the simulated probability density of rzAUC under H0 in the comparison of two balanced classes. The density of rzAUC distribution seems to approach a normal probability function when the sample size is increased. However, in the presence of a high number of samples the distribution tends to be slightly leptokurtic and negative skewed (Figure 1S, panel (f)).
Expected value of rzAUC as a function of the sample size under the null hypothesis
n0 = number of controls; n1 = number of cases.
Estimates obtained by 105 random samples.
Variance of rzAUC as a function of the sample size under the null hypothesis
n0 = number of controls; n1 = number of cases.
Estimates obtained by 105 random samples.
3 Estimates of statistical power of rzAUC
Proportion of positive results (%) for test on rzAUC and AUC under different TMs distribution assumption
TMs: tumour markers.
Asy = asymptotic normal test; Per = permutation test (based on 2000 random permutations); n0 = number of controls; n1 = number of cases; and *H0 is true.
Δμ: expected difference in TM means between cases and controls.
3.1 Binormal model with equal variances
In the case of binormal model with equal variances, when the means in the two groups were equal (difference between the two means, Δμ = 0) the proportion of positive results represented an estimate of the test bias under H0, whereas in the presence of a positive Δμ value, such a proportion provided an estimate of the corresponding statistical power. Both asymptotic and permutation tests on rzAUC showed a rather good agreement with the expected α-value (0.05) under H0 and with the corresponding results from the standard analysis on AUC, even if the asymptotic test on rzAUC was slightly biased in the presence of small sample size (Table 3). Under H1 (Δμ from 0.5 to 2.0) a good agreement between the asymptotic and the permutation analysis was observed for both statistics at any sample size, except for a small loss of power for the asymptotic test for rzAUC, which tended to decrease when increasing Δμ. The test for AUC showed the highest statistical power, but the difference between the two tests tended to disappear with increasing both the sample size and Δμ.
3.2 Binormal model with different variances
The estimate of statistical power of the two tests under a binormal model with different variances was obtained setting to 1.0 the variance among controls and to 2.0 that among the cases. Results from permutation and asymptotic methods were in a rather good agreement for both tests in any analysis, except for a small loss of statistical power in the permutation test for rzAUC either at sample size ≤20 or when the difference between means (Δμ) was lower than 1.5 (Table 3). The test for rzAUC showed a higher statistical power than that for AUC at Δμ = 0.5, while for larger Δμ values the two tests had a comparable performance. At Δμ = 0, the proportion of positive results was higher for the test for rzAUC (both from permutation than asymptotic analysis) than that for AUC. However, it should be noted that in this situation results for the new proposed test do not estimate the test bias under H0, because the corresponding ROC curve crosses the chance line.
3.3 Normal-binormal model
Simulations for the normal–binormal model (corresponding to ROC curves like that in Figure 3(c)) were carried out setting to 1.0 the variance for any distribution (i.e. two normal distributions among the cases and one normal among the controls). Mean values were set to zero for both the first subgroup of cases and the class of controls, while it ranged from 0.0 to 2.0 for the second subgroup of cases. A sampling ratio of 2:1 was adopted for the two subgroups of cases (the largest one was that with the same distribution of the controls). In this analysis, Δμ represents the difference between the means between the two groups of cases or (equivalently) between the second group of cases and the controls.
Permutation test showed a higher statistical power than the asymptotic test, but this difference tended to disappear with increasing the sample size (Table 3).
3.4 Binormal–binormal model
Under the binormal–binormal model, in the two subgroup of cases and controls with the same mean, variance was 1.0, the mean value was 0.0 and the sample size was the same. In the remaining two subgroups of controls and cases, mean values were 1.0 for the controls and were left to vary from 1.0 to 3.0 in the subgroup of cases (in Table 3 such differences were expressed as distances Δμ between such subgroups, which accordingly varied from 0.0 to 2.0). Variance and sample size were equal to those of the former two subgroups.
In the presence of small sample size, the normal test for rzAUC showed a loss of statistical power compared with the permutation approach. Tests for rzAUC outperformed that for AUC, with differences becoming more pronounced when increasing Δμ values.
3.5 Zero-inflated binormal model
For the zero-inflated binormal model the same sample size was adopted for the two subgroups of cases and controls with mass at zero and for the remaining two subgroups. For these latter, a binormal model with equal variances (σ2 = 1.0) was used, with a mean value of 1.0 in the subgroup of controls and a mean value varying from 1.0 to 3.0 in the subgroup of cases (corresponding to Δμ from 0.0 to 2.0 in Table 3). Negative values extracted from the normal distributions were replaced with zero. At Δμ = 0.0, the test for AUC showed a proportion of false positive tests close to 0. The asymptotic test for rzAUC was also biased, especially for low sample size, while for the permutation test the false positive proportion was quite similar to the selected α-value (0.05).
The permutation approach outperformed the asymptotic test for rzAUC in any analysis. Both tests for rzAUC showed a higher statistical power than that for AUC.
In summary, results of power analysis reported in Table 3 suggest that, in most cases the asymptotic test for rzAUC seems to be appropriated when at least 20 samples are present in each class, whereas for smaller sample sizes the permutation test should be preferred.
4 Choosing between complete and rROC curves
Distribution of the number of samples j excluded by rROC analysis at the conventional 0.05 α level for not informative (Δμ = 0) and for proper ROC curves (Δμ > 0)
rROC: restricted receiver operating characteristic.
n0 = number of controls; n1 = number of cases; % = percentages of rROC curves selected; MDP = mean difference percent between rAUC and AUC; and n.e. = not evaluable.
Δμ = expected difference in TM means between cases and controls.
Simulations included only proper ROC curves, because the standard ROC analysis approach should be preferred in the case of a proper ROC curve, due to the highest statistical power of the test for AUC. Moreover, the case of non-informative ROC curve was also considered. The distribution of j-values was obtained from 1000 simulated TMs. The corresponding statistical significance of rzAUC was estimated by the test based on the normal approximation (equation (6)), except when the sample size was lower than 20 to prevent the loss of statistical power. In that case, a test based on 2000 random permutations was employed.
Table 4 reports the percentages of rROC curves that would have been selected using the above cited three criteria and some corresponding percentiles of the j statistic, which corresponds to the number of samples excluded from the curve. The first column corresponds to the results of rROC analysis only (case a). The number of excluded samples was about 5% under H0 and tended to increase with increasing Δμ, while the corresponding median value of j tended to decrease. In particular, for Δμ ≥ 1.0 the median value of j was 0 in any comparison and for Δμ ≥ 1.5 also the 75° percentile equals 0 in all comparisons except one, indicating that the rROC analysis tends to provide results very similar to those from the standard ROC approach when the underline TM distributions are well separated.
The second column in Table 4 (case b) shows the results of the combination of restricted and standard ROC analysis based on the comparison between the corresponding p-values. The proportion of selected rROC was slightly lower than those observed by the rROC only (first column) at Δμ = 0.0, while the median value of j was slightly higher. Proportion of errors tended to increase with increasing Δμ up to Δμ = 1.0 and to decrease subsequently. Finally, when only rROC curves corresponding to DMP >20% were selected (case c, third column in Table 4) the proportion of errors clearly decreased ranging between 4% and 7% for Δμ ≤ 1.0 and dropping to 0–2% for Δμ > 1.5, while the corresponding j-values decreased accordingly.
Taken together, these results suggest that in the presence of a statistically significant test for the rROC curve and a MDP > 20% standard ROC analysis may be replaced by rROC ones with an acceptable proportion of errors (about 5% or less).
5 Combining information from restricted and standard ROC curves
The main limit of the rROC analysis is the exclusion from the ROC curve of the j subjects whose TM values are lower than Cj, the cut-off that separates informative from non-informative values (as illustrated in Section 2). To recover this drawback, information from two or more TMs can be combined. A new approach will be illustrated, which extends the traditional method of applying a logistic regression model to two or more marker values in order to obtain a risk score (RS) for classification purposes. 19 RS is usually employed as a new TM in a standard ROC analysis and a binary test is identified from a corresponding suitable cut-off.
In a standard approach, a logistic regression model with main effects can be employed, and RS estimated as a weighted sum of the TMs value (TMi), by adopting as weights the corresponding regression coefficients βk
In the presence of rROC curves, the corresponding markers can be introduced in the regression model using a nested approach
In the further section an application of the proposed method will be illustrated using a real data set (Section 6.3).
6 Application of the new method to a real data set
6.1 Description of the data set
The new method of ROC curve analysis based on rROC was applied to a real data set that collected data of peripheral blood concentrations of seven putative TMs from 15 patients affected by neuroblastoma and from 20 healthy controls, namely: tyrosine hydroxylase (TH), beta-1,4-N-acetyl-galactosaminyl transferase 1 (GD2), dopamine-decarboxylase (DDC), doublecortin (DCX), embryonic lethal, abnormal vision-4 (ELAV-4), sialyltranferase ST8SiaII (STX) and paired-like homeobox 2b (Phox2b). 20 Neuroblastoma is the most frequent extra cranial solid tumour of childhood, presenting in localized (stages 1, 2 and 3) and metastatic (stages 4 and 4S) forms. 21 Since neuroblastoma originates from neuronal cell precursors, the presence of neuronal specific RNAs in peripheral blood samples have been proposed as a measure of tumour cell contamination. However, low expression of neuronal RNAs from peripheral blood cells have been documented also in samples from healthy subjects. 22 Patients with localized neuroblastoma are supposed to have little, if any, tumour cells in blood, but their presence may have important prognostic value. 21 Therefore, it is of the outmost importance to identify the TMs able to discriminate patients with tumour infiltration in the blood from those without.
Concentration values a of six putative neuroblastoma TMs measured in the study subjects
TH: tyrosine hydroxylase; GD2: beta-1,4-N-acetyl-galactosaminyl transferase 1; DDC: dopamine-decarboxylase; STX: sialyltranferase ST8SiaII; DCX: doublecortin; and Elav: embryonic lethal, abnormal vision.
Values were obtained by applying the ΔΔCq formula to the readout of real-time PCR equipment (see text for details).
n.a. = non-applicable.
6.2 Comparison between standard ROC and rROC analysis
Figure 4 (panels (a) to (f)) shows the ROC curves corresponding to the TM values in Table 5. In each panel, results of both standard ROC analysis (AUC, zAUC) and rROC analysis (rAUC and rzAUC) are reported with the corresponding p-values obtained by the normal tests (pNorm) and by 2000 random permutations (pPermut). Moreover, in each plot the cut-off of highest accuracy J and the cut-off C, which identifies the rROC according to equation (5), are also reported with the corresponding 95% confidence intervals (in brackets) estimated by the percentiles of the bootstrapped distribution obtained from 5000 bootstrapped samples.
24
ROC curves corresponding to the TM values in Table 3. In each panel results of both standard ROC analysis and rROC analysis are reported with the corresponding p-values obtained both by the normal tests (pNorm) and by 2000 random permutations (pPermut). rROC: restricted receiver operating characteristic; TM: tumour marker. FPF: false positive fraction; TPF: true positive fraction.
All the considered markers corresponded to a non-proper positive skewed ROC curve. In particular, TH, DDC, STX and DCX (Figure 4, panels (a), (c), (d) and (e)) showed a behaviour similar to that of the theoretical ROC curve in Figure 2(d), due to the presence of numerous null values in the TM distribution (Table 5). The shape of the curve corresponding to Elav (panel (f)) is due to the presence of a range of small positive values of marker concentration, similarly distributed among the two classes under study and it is consistent with a behaviour like those described by the ROC curves in Figure 2, panels (b) and (c). Finally, the ROC curve for GD2 (panel (b)) showed an intermediate behaviour, due to the presence of a range of small non-informative TM concentrations that also included some zero values. Results from permutation tests were consistent with those from normal tests for both standard ROC and rROC analysis. Using a conventional 0.05 level of statistical significance (one sided test), STX and DCX would have been rejected as TM by the standard ROC analysis (AUC = 0.64, pNorm = 0.086 for STX and AUC = 0.63, pNorm = 0.103 for DCX, respectively), whereas the test for rROC was statistically significant in both cases (pNorm = 0.045 and pNorm < 0.001, respectively). For the remaining curves both approaches provided evidence of a possible application of the corresponding TM for the diagnosis of localized neuroblastoma. Interestingly, the two approaches provided the same result (AUC = rAUC) for DDC, whose ROC curve (Figure 11S, panel (c)) was rather similar to a proper ROC curve (Figure 1, panel (a)). All TMs, except DDC, showed a MDP > 20%, suggesting that the application of the rROC analysis may be appropriated (Section 4).
Results of standard ROC and rROC analysis applied to data in Table 5 and ROC curves in Figure 4
rROC: restricted receiver operating characteristic; TM: tumour marker; TH: tyrosine hydroxylase; GD2: beta-1,4-N-acetyl-galactosaminyl transferase 1; DDC: dopamine-decarboxylase; STX: sialyltranferase ST8SiaII; DCX: doublecortin; and Elav: embryonic lethal, abnormal vision.
Nexcl = total number of values excluded from rROC; Ncases = number of values among cases excluded from the rROC; and Ncontr = number of values among controls excluded from the rROC.
6.3 Identification of a classification rule based on the RS method
Results of the classification using a RS obtained by the combination of couple of the tumour markers in Table 3 by either a standard approach or a nested model based on the rROC analysis
RS: risk score; rROC: restricted receiver operating characteristic; TMs: tumour markers; TH: tyrosine hydroxylase; GD2: beta-1,4-N-acetyl-galactosaminyl transferase 1; DDC: dopamine-decarboxylase; STX: sialyltranferase ST8SiaII; DCX: doublecortin; and Elav: embryonic lethal, abnormal vision.
Legend: Se = sensitivity; Sp = specificity; Acc = diagnostic accuracy; Global validation: the best couple of markers (based on the highest AUC for the RS) selected at each step from the whole data set.
Convergence was not achieved.
The new method based on rROC analysis and regression models with nested variables clearly outperformed the standard one showing an equal to higher accuracy in each comparison. However, in two cases (corresponding to the inclusion in the model of GD2 and DCX, and DCX and Elav, respectively) the convergence was not achieved.
In cross-validation the new method outperformed the standard one in five cases, had a lower accuracy in four and an equal accuracy in the remaining six comparisons. In the global cross-validation, the accuracy of the new method barely exceeded that of the standard one (88.6% vs 85.7%).
7 Discussion
rROC curves represent new simple tools for the analysis of TMs for diagnostic purposes. In this article a new statistical test, associated to the area under a rROC curve, has been proposed. In the presence of positive skewed ROC curves, the new test outperformed the standard approach both on simulated distributions and on a set of real TMs values. In particular, in the real data set two TMs, namely STX and DCX, would have been treated as non-informative by standard ROC analysis.
Standard measures of summary accuracy like AUC or the Youden’s index were demonstrated to be a quite simple, powerful and useful tools for diagnostic purposes in many medical fields.2,3,8 However, when applied to not proper ROC curves, standard ROC analysis may be unable to extract most of the relevant information or it may even provide unreliable results. 25 In Clinical Epidemiology not proper ROC curves are often encountered, especially in the analysis of putative TMs. In particular, positive skewed ROCs may results from the presence of similar values at low TM concentrations in the two classes under study. Many reasons may be advocated for this behaviour, including the presence of a detection threshold in the measure device, the presence of a subgroup of cases with a TM expression similar to that of controls, and the presence among controls of individuals with a rather high expression of TM in the absence of neoplastic cells, as briefly illustrated in the Introduction and shown by some examples in Figures 2 and 3. With respect to the real data set analysed in this article, 20 the shape of the corresponding ROC curves reported in Figure 4 indicates that a mix of these different situations have occurred. In particular, zero-inflated distributions may have arisen from the detection threshold of the RT-qPCR method that is unable to perform more than 40 amplification cycles, while bimodal distributions among cases may be related to the fact that all the considered TMs can be expressed in some proportion also by normal cells as a consequence of the so called illegitimate transcription. 23 Furthermore, some patients with localized neuroblastoma, which have little tendency to spread to distant tissues, including blood, may have TMs concentration values similar to those of healthy individuals.
Information from the shape of a ROC curve seems to be at least as important as information from summary indices of accuracy like AUC. In recent years, modern ROC analysis has tried to develop new approaches to extract and combine all available information from a TM, exploring the ROC regions with informative values. 26 Probably the most important and reliable method is based on pAUC, which was demonstrated to be a useful tool for diagnostic purposes in the presence of asymmetric ROC plots.13,17,26 The method based on rROC, described in this article, may be considered as an extension of the methods based on pAUC. In fact, the area under a rROC curve (rAUC) is equivalent to the pAUC identified by the cut-off C, divided by the corresponding 1-specificity value. However, there are two major differences with the pAUC based approaches, namely: (a) the cut-off that identifies rROC is selected by a recursive procedure coupled with an index of accuracy (rzAUCj, equation (3)) and (b) samples corresponding to the non-informative TM concentrations are considered as missing values and not as negative test. For this reason at least another diagnostic marker, if available, is needed in order to complete the classification analysis, thus allowing those subjects, whose TM levels lie in the range of ‘not informative values’ (TM concentrations <C), to be allocated in either one class under study. Using the real data set of neuroblastoma markers, 20 a combination of information from both standard and rROC analysis via a logistic regression model provided a classification rule based on the corresponding RS whose accuracy was higher than that of any other combination of couples of TMs. However, a more reliable estimate should have been obtained from an external validation set, which unfortunately was not available. The small sample size of the data set (20 controls and 15 cases) just allowed the application of the leave-one out cross-validation. Most results from this analysis indicated a better performance of the new proposed method, but accuracy of the global validation was only marginally higher than that of the standard analysis. Furthermore, the use of regression models with nested variables in some instances provided unstable estimates due to a lack of convergence even in the analysis of the whole data set, probably as a consequence of the small sample size. Finally, nested models may be more prone to overfitting than the corresponding main effect models due to the inclusion of a higher number of predictors. Further analyses, based on both extensive simulations and real databases of TMs with different sample sizes, are needed to estimate the accuracy and the precision of the new proposed method and to assess the advantage of using rROC analysis in classification based on RS from a regression model.
Another limitation of the proposed method is that the statistical properties of the new proposed test were explored only via extensive simulations, whereas an analytical formula for rzAUC distribution under H0 is not available. This limit might make difficult its application to very large data sets, such as multicentre follow up study, where the analysis by random permutations is rather time-consuming. Furthermore, only left restriction has been considered in this article in the presence of positive skewed ROC curves, whereas in clinical setting many other types of not proper curves may be encountered. In particular, applying a right restriction, the proposed method should be extended to negative skewed ROC, that may originate by a ceiling detection threshold in a diagnostic device or by the presence of some spiked values among the controls. However, the extension of the rROC method to include either type of restriction will require a very large number of simulations and the use of new databases of real data, and is behind the scope of this investigation. Finally, the method should also be extended to allow the inclusion of an utility function 16 in the evaluation of optimal cut-offs, since false positive and false negative errors may have different costs in different situations. Nonetheless, results from this investigation indicate that rROC curves may be a new simple and useful tool to explore the diagnostic utility of putative TMs. In the coming future, new statistical tests should be developed to extract information from different kinds of TM distributions, possibly exploiting the properties of the corresponding rROC plots. The development of a comprehensive approach to combine any relevant information from both standard and modern ROC analysis will provide an optimal framework for the evaluation of the diagnostic potential of new TMs.
Footnotes
Acknowledgements
The authors thank Dr Paolo Bruzzi (National Cancer Research Institute of Genoa) for providing precious advice.
Funding
This study was partly supported by the Ligurian Region and by the Italian Neuroblastoma Foundation (Fondazione Italiana per la Lotta al Neuroblastoma). B.C. is a recipient of a grant from the Italian Neuroblastoma Foundation.
Declaration of conflicting interest
None declared.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
