Abstract
It has become easy to obtain multivariate chemical data of high dimensions. However, it may be expensive or time consuming to obtain a large number of samples or to acquire reference measures, so the number of samples available for multivariate calibration modelling may be limited. If data contains nonlinear relationships, nonlinear methods are required for the calibration task. The combination of limited amounts of data of high dimensions and highly flexible nonlinear methods may result in overfitted models which in turn perform badly on new data. Therefore, for real world applications, it is desirable to understand how the sample size affects model prediction performance. For this purpose, we compared partial least squares regression, artificial neural network, and support vector regression applied to three real world nonlinear datasets of which two were of high dimensions. We evaluated the effect of calibration sample size (i) on test set performance, including variation in test set performance due to sampling variation and (ii) tested if the cross-validated performance was adequate for assessing the predictive ability. We demonstrated the applicability of artificial neural network and support vector regression for real world data of limited size and showed that support vector regression had advantages over artificial neural network: (i) fewer calibration samples were required to obtain a desired model performance, (ii) support vector regression was less sensitive to sampling variation for small sample sets and (iii) cross-validation was an approximately unbiased option for evaluating the true support vector regression model performance even for small sample sets.
Keywords
Introduction
During the last decades, there has been an increase in sophisticated chemical analytical methods generating data with a very high number of variables. Multivariate modelling methods are usually required to relate such data to the desired biological, chemical or physical property. In particular, partial least squares regression (PLS) is widely applied for chemical data analysis. However, assuming linearity sometimes leads to suboptimal models, because the relationships modelled are indeed nonlinear. Near infrared (NIR) spectra and quantitative structure–activity relationship (QSAR) data are examples of data often exhibiting various types of nonlinearities. Changes in the physical and chemical constitution of a sample, such as temperature and pH, cause spectral shifts or in other ways non-fulfilment of Beer’s law.1–3 In the case of QSAR data, the descriptors may be intrinsically nonlinearly related to the property of interest.
Nonlinear methods are not widely applied for modelling of chemical data. The reasons for avoiding these methods are many and include too little knowledge on (i) which method to choose, (ii) how to optimize the model including selection of meta-parameters, (iii) sample size requirement, (iv) how to deal with noise and outliers and (v) computer power requirement. Dogmatic perceptions such as common practice and possibly faulty views of nonlinear methods being very difficult to deal with obstruct the dissemination of these methods. In order to gain wider currency, information on how to build and interpret nonlinear models should be easily available.
In this paper, we address the prediction performance of calibration models PLS, artificial neural network (ANN) and support vector regression (SVR) under different calibration sample sizes. This is important because although it has become easy to measure a high number of variables for many types of analyses, it may still be expensive and time consuming to obtain a large number of samples or to acquire adequate reference measures for model development. Therefore, for real world applications, it is desirable to understand how the sample size affects model prediction performance.
Nonlinear modelling of limited amounts of data is particular challenging when data are high dimensional and this is partly due to the curse of dimensionality where model flexibility increases with increasing data dimensionality. In general, a more flexible model needs more calibration data to determine its parameters accurately. Nonlinear methods add to the curse of dimensionality, since they intrinsically offer more flexibility than linear methods do, simply because they can fit more complicated relations.
SVR is a nonlinear method which is characterized by high generalization ability. 4 SVR and support vector machines (SVM) for classification entered the chemical literature around 2000.5–9 Since then, SVR has gained increasing interest within chemometrics,10–12 but its dissemination in real world applications is still limited compared to ANN,10,13–15 where the latter has been a tool for analysis of chemical data since late 1980s. 16 ANN is generally considered to require a very large number of calibration samples to avoid overfitting.17–19 Numerous papers compared the performance of SVR and ANN on various types of chemical data, mostly to the benefit of SVR.11,19,20 However, only few examples of a systematic investigation of the impact of calibration sample size on prediction performance exist for regression21,22 and for classification.23,24 All obtained similar results, namely that SVM and SVR performed better than ANN for all sample sizes but especially for small sample sizes.
None of these previously examined datasets were of very high dimensions. This paper reports a systematic study of the effect of sample size on model performance for high dimensional chemical data. For this purpose, PLS, ANN and SVR were applied to (i) a dataset of NIR spectra measured at 450 wavelengths, predicting sucrose. To further exemplify and understand the performance, the models were also investigated on (ii) a QSAR dataset containing 150 2D descriptors, predicting melting point, and (iii) a dataset containing eight quantitative attributes of concrete, predicting compressive strength. The obtained models were evaluated as follows: (i) model performance was evaluated as root mean squared error of an independent test set (RMSEP) and the effect of sampling variation on model performance was evaluated as the variation in RMSEP as a result of models built on 30 realizations of a calibration dataset of a defined number of samples (n). Moreover, (ii) the adequateness of using cross-validation for assessing predictive ability was evaluated as RMSEP relative to the root mean squared error of cross-validation (RMSECV): percentage of (RMSEP–RMSECV)/RMSECV which we named cross-validation optimism (CV optimism). A low CV optimism would reveal RMSECV as a reliable measure of the true model performance and hence could be used for model validation. This in turn could reduce the number of samples required for a calibration task, since setting aside samples for an independent test set could be avoided.
Theory
Support vector regression
Overview of main characteristics of PLS, ANN and SVR.
ANN: artificial neural network; PLS: partial least squares regression; SVR: support vector regression.
Further, in order to obtain a model which generalizes beyond the training data, a penalty is put on large regression coefficients through a penalty on the length of
The slack variables ξi, ξi* are introduced for the situations where f(
Here
One of the most widely used kernel functions is the radial basis function (RBF) kernel
7
The kernel width, σ, controls the size of a neighbourhood around a support vector where it influences predictions of new samples.
The three meta-parameters C, ɛ and σ must be selected by the user prior to model estimation. As they jointly influence model generalisation, they should be optimized simultaneously.
PLS and ANN
PLS and ANN will not be described here, but instead the reader is referred to the literature.27–29 The main features, though, are outlined in Table 1 along with the characteristics of SVR. The linearity of PLS imposes limitations on the model but PLS only requires determination of a single meta-parameter, namely the number of latent variables. It is deterministic and the solution is unique under most circumstances. The nonlinear ANN requires optimization of a number of meta-parameters including selection of transfer function(s) and number of optimization cycles. Further, the ANN solution is not unique and depends on initialization. Compression of data e.g. by PLS prior to modelling is often used to bound the model flexibility.
Data descriptions and methods
Overview of main characteristics of the three datasets.
NIR: near infrared; QSAR: quantitative structure–activity relationship.
Quantitative attributes of concrete.
NIR spectra from a sugar factory for prediction of sucrose
The main dataset was from a sugar factory and contained NIR measurements from four process steps aimed for prediction of sucrose. 30 The dataset contained nonlinear effects due to changes in the physical and chemical constitution of the process stream during production. A total of 1800 samples were measured at 400–1890 nm with an increment of 2 nm. Models were built on 450 selected variables. 30 Sucrose, the reference variable was measured in weight percentage. The dataset was split into calibration (66%) and test (33%) set using the so-called onion method as implemented in PLS_Toolbox version 7.9 (Eigenvector Research, Inc. Wenatchee, WA). 31 The splitting was manually evaluated in order to ensure that the two sets seemed to be similarly distributed.
QSAR data for prediction of melting points
Prediction of melting points are important for various applications such as assessment of hazardousness of chemical compounds in REACH a or in the medicinal and environmental chemistry for estimating solubility. 32 Prediction of melting point from QSAR data is known to be difficult, as indicated by high prediction errors of previously reported models (RMSEP >31℃).32,33 The present QSAR dataset has a total of 12,589 compounds with 150 two-dimensional descriptors and originated from several different databases. 33 The dataset was split into calibration and test set similar to the NIR dataset.
Concrete data for prediction of compressive strength
The concrete dataset was included in the study as a third and very different dataset in order to investigate if the sample size dependence was mainly data type dependent or if a more general tendency existed. This dataset was an extended version of the dataset used by Yeh. 34 It consisted of eight quantitative attributes of manufactured concrete for a total of 1030 samples for prediction of the compressive strength of high performance concrete. According to Yeh 34 the compressive strength of concrete was a highly nonlinear function of age and ingredients. The eight attributes showed very little linear correlation to the compressive strength. The dataset was split into calibration (83%) and test set (17%) using the onion method.
Study design
The following procedure was applied to all three datasets: New datasets of decreasing number of calibration samples were obtained by sampling 50% of the previous calibration set. This was done until the obtained training set reached a size of approximately 18 samples. Thirty series of such dataset realizations were generated.
For generalisation of the results in this study it was assumed that the sampling of calibration data was drawn from an infinite pool of samples, which directly implied independence between calibration data sets. However, the data were real and consequently from a finite pool of samples, which introduced dependency between calibration data sets. For the calibration sample sizes close to the size of the total sample pool this entailed that the variation in RMSEP was biased downwards due to the calibration models being constructed from data with overlapping samples. This leads to a 40% and 15% underestimation of the variation in RMSEP of models based on 50% and 25% of the total sample pool, respectively. For the remaining sample sizes, the effect was negligible. Moreover, it did not affect the comparison between the different methods, which is why we have not corrected the results in this regard.
Methods and software
The NIR data were corrected using a standard normal variate (SNV) followed by ordinary mean centring of the columns of the data matrix. The QSAR and concrete datasets contained discrete x-variables and hence autoscaling was applied. All reference variables were mean centred. For each data type, strongly outlying samples were removed from an initial PLS model built on all calibration samples. Automated selection of meta-parameters was applied to all methods selecting the setting with the smallest error from a 10-fold cross-validation (RMSECV). Models were validated by RMSEP of the independent test set. Furthermore, CV optimism was measured as the percentage (pct.) of (RMSEP–RMSECV)/RMSECV.
PLS
One to 20 latent variables were investigated for the minimum RMSECV in case of the NIR and QSAR data and a maximum of eight latent variables were investigated in case of the concrete dataset.
ANN
A feed forward, backpropagation neural network was used. To prevent overfitting, it comprised a single hidden layer and data were compressed by PLS in order to reduce the number of nodes in the input layer. A grid-search was used to select the optimal number of latent variables in the PLS model and optimal number of nodes in the hidden layer. The maximum values of the meta-parameters were: NIR data: nine latent variables and nine nodes, QSAR data: six latent variables and eight nodes, concrete data: eight latent variables and 12 nodes. A maximum of 20 learning cycles were used in the model optimization, and the actual number of cycles, number of latent variables and nodes was selected by the minimum of RMSECV. The adopted ANN algorithm used random initialization of the weights and there was no weight decay for the single hidden layer.
SVR: Grid search was applied for selection of C, ɛ and the σ parameter in the RBF kernel. All three meta-parameters were searched in a log2 scale and selected by cross-validation.
Calculations were made in MATLAB R2014a (The MathWorks, Inc., Natick, MA) and PLS_Toolbox version 7.9 (Eigenvector Research, Inc. Wenatchee, WA).
Results and discussion
NIR data
Figure 1(a) shows box plots of the test set error of the models built on the NIR spectra. All modelling methods showed similar influence of sample size on model performance: The test set error increased very little when the number of calibration samples was reduced from 1134 to 143. A further reduction in sample size lead to a steady increase in test set error for PLS and SVR and an exponential increase for ANN. However, the increases were moderate in general.
NIR data. PLS (blue), ANN (green) and SVR models (red). (a) Box plot of RMSEP (%w/w). (b) Box plot of CV optimism expressed as pct. of (RMSEP − RMSECV)/RMSECV. (a) and (b) x-axis: number of samples in calibration set. For each method one model was developed for n = 1134, 30 models were developed for the remaining sample sizes. Outliers are shown as circles and defined as observations larger than q3 + w(q3 − q1) or smaller than q1 − w(q3 − q1) where q1 and q3 refers to the lower and upper quartile, respectively, and w = 1.5 and w = 3.0 are white and black, respectively. ANN: artificial neural network; CV: cross-validation; NIR: near infrared; PLS: partial least squares regression; RMSECV: root mean squared error of cross-validation; RMSEP: root mean squared error of an independent test set; SVR: support vector regression.
The almost constant performance of all methods for large n suggested that all calibration models had approached their minimal model error. The RMSEP of SVR was close to the reference uncertainty of 0.5 30 indicating that SVR had sufficient flexibility to model the true relationship in the data. The higher error of PLS was likely due to its intrinsic limitations in modelling nonlinearities.
SVR showed the smallest median RMSEP in the entire range of calibration sample sizes. For n ≥ 143 samples, one could be almost certain to obtain a SVR model that performed better than a corresponding PLS or ANN model as shown by no or very little overlap in the distribution of the obtained test set errors of the methods. For the smallest sample sizes, test set error was similar for all methods except ANN which showed a sudden increase in RMSEP variation for 19 calibration samples. Yet, ANN performed surprisingly well even for small sample sizes.
It was unexpected that PLS would perform better than ANN for large sample sizes, in part because of the automated selection of PLS latent variables, which selected a, from a chemical viewpoint, unrealistically high numbers of variables. E.g. all models based on 567 calibration samples contained the maximum of 20 latent variables. b However, these models that contained a high number of latent variables did perform well.
Moderate increases in RMSEP variability were observed for all methods when the number of calibration samples was reduced, demonstrating that one should be increasingly concerned that the calibration samples indeed represented future samples well when the number of calibration samples was limited. This was especially an issue for ANN models built on 19 samples.
PLS and SVR did not show much CV optimism in general until n ≤ 37 samples and the variability only increased slowly, illustrating highly reliable RMSECV estimates (Figure 1(b)). In contrast, ANN displayed a consistently positive central distribution (CD – the interval from the lower 25% to the upper 75% quartile depicted as the box in the box plot) of the CV optimism, which was slowly growing until 37 calibration samples. After that, the CV optimism, as well as the RMSEP, increased significantly, pointing at a breakdown in ANN generalization ability.
QSAR data
Figure 2(a) outlines the test set performance of models built on the QSAR data. Generally, the test set errors were high which agreed with previously reported models.32,33,35 Possibly the high error derived from a deficiency in the descriptors for explaining the melting point.
35
Moderate overall reduction in ANN and SVR model performance was observed, as the median RMSEP increased by 75% for ANN and 81% for SVR when the number of calibration samples was reduced from 9964 to 20. In contrast, the median RMSEP of PLS increased by 537%.
QSAR data, PLS (blue), ANN (green) and SVR models (red). (a) Box plot of RMSEP (℃). (b) Box plot of CV optimism expressed as pct. of (RMSEP − RMSECV)/RMSECV. (a) and (b) lower plot is zoom in of the upper plots. x-axis: number of samples in calibration set. For each method one model was developed for n = 9964, 30 models were developed for the remaining sample sizes. Outliers are shown as circles and defined as observations larger than q3 + w(q3 − q1) or smaller than q1 − w(q3 − q1) where q1 and q3 refer to the lower and upper quartile, respectively, and w = 1.5 and w = 3.0 are white and black, respectively. ANN: artificial neural network; CV: cross-validation; PLS: partial least squares regression; QSAR: quantitative structure–activity relationship; RMSECV: root mean squared error of cross-validation; RMSEP: root mean squared error of an independent test set; SVR: support vector regression.
SVR performed significantly better than PLS and ANN. Indeed, there was no overlap of PLS and SVR CD for any calibration sample size. There was no overlap between ANN and SVR CD for five sample sizes, but they performed similarly for 623 and 312 calibration samples.
The median RMSEP of SVR models increased steadily throughout the entire range of calibration samples tested. PLS and ANN median RMSEP were fairly flat for large n but showed a growth increase for n = 1246 to 20, with faster growth for PLS than for ANN. Hence, SVR was less affected by small calibration samples sizes than PLS and ANN. The levelling off in PLS and ANN RMSEP for large n indicated that they approached the maximum capacity of the methods for modelling the data. The levelling off appeared for similar sample sizes, but ANN obtained lower test set error than PLS, indicating its better ability for modelling nonlinear relationships. On the other hand, the steady improvement in SVR RMSEP indicated that SVR could continue to improve significantly if built on more calibration samples than included in this study.
SVR was highly robust against sampling variation as indicated by the generally low RMSEP variation as well as its limited increase for decreasing number of calibration samples. PLS was not affected much by sampling variation until n = 156, at which point some models obtained test set errors several times larger than the median error. Reducing the number of calibration samples further led to extreme test set performances for an increasing number of models. A similar pattern was observed for CV optimism of the PLS models (Figure 2(b)). These patterns were in stark contrast to ANN and SVR, although they were all presented with the exact same dataset realizations. Therefore, these observations pointed at PLS being insufficient in modelling the true structure in the data, likely due to linearity imposing limitations upon model flexibility. In combination with the high level of PLS model error in general and low sample size, it made PLS highly sensitive to sampling variability. The nonlinear methods, on the other hand, apparently matched the data structure better and had lower error level, so they could more accurately aim for the relevant variation in data regardless of the specific data representation. However, ANN was more sensitive than SVR to sampling variation revealed by consistently higher variation in test set error.
Figure 2(b) shows a significant increase in ANN CV optimism reaching 380% on average (median CV optimism) for n = 20 samples. This indicated that the models fitted the calibration data very (too) well and hence generalized poorly. However, the test set performance was not extreme compared to SVR. Hence, our results demonstrated that it was possible to obtain fairly well performing ANN models although they overfitted the calibration data. It was crucial, though, to use an independent test set for model validation. The CV optimism of SVR was low for all sample sizes, why RMSECV was a reliable measure of the true model performance even for as little as 20 calibration samples.
Concrete data
SVR and ANN had very overlapping test set error of the concrete dataset in general. However, an increasingly smaller test set error of SVR compared to ANN was observed for n = 53 to 14 samples. Both nonlinear methods continued to improve for high n, whereas PLS started to level off at n = 53. That is, PLS approached its minimum error at a relatively small sample size whereas no indication of a similar approach was observed for ANN and SVR for the sample sizes studied.
Small differences in RMSEP variability between the methods indicated that PLS was least sensitive to sampling variation, which agreed with the expectation that a less flexible model will over fit less and hence be less sensitive to small variations in calibration data.
The concrete models showed CV optimism profiles similar to the NIR models. However, for n = 14, all methods showed similarly CV optimisms, meaning that RMSECV became an equally uncertain measure of model performance (Figure 3(b)).
Concrete data. PLS (blue), ANN (green) and SVR models (red). (a) Box plot of RMSEP (MPa). (b) Box plot of CV optimism expressed as pct. of (RMSEP − RMSECV)/RMSECV. (a) and (b) x-axis: number of samples in calibration set. For each method one model was developed for n = 847, 30 models were developed for the remaining sample sizes. Outliers are shown as circles and defined as observations larger than q3 + w(q3 − q1) or smaller than q1 − w(q3 − q1) where q1 and q3 refer to the lower and upper quartile, respectively and w = 1.5 and w = 3.0 are white and black, respectively. ANN: artificial neural network; CV: cross-validation; PLS: partial least squares regression; RMSECV: root mean squared error of cross-validation; RMSEP: root mean squared error of an independent test set, SVR: support vector regression.
Comparison of the methods across the datasets
Model performance
All methods showed a similar overall performance profile for decreasing amounts of calibration data across data types; the model performance decreased when the amount of calibration data was reduced. For a specific method and specific dataset, the calibration sample size dictated the model performance; an observation supported by classical statistical theory.
In the present study, SVR performed better than PLS and ANN for all sample sizes of all datasets, with only a few exceptions where PLS or ANN performed similar to SVR. Hence, our results demonstrated the high generalization performance of SVR and indicated that it applies to large as well as small sample sizes. For large sample sets it was noteworthy that SVR performed better than the nonlinear alternative, ANN, for two out of three datasets.
PLS was least affected by sample size, except for the QSAR data. But contrary to our expectations, this was due to poor performance on large sample sets only, and not due to good performance on small sample sets. In general, only moderate reduction in test set performance was observed (a factor of three for median RMSEP) when the calibration data were reduced from several thousand samples to approximately 20 samples. This was particularly surprising for ANN, which we expected to display a strong decrease in model performance when sample sizes were reduced. Hence, our results challenge the dogmatic perception of ANN requiring a high number of calibration samples in order to perform well. However, the ANN modelling was not conducted directly on the raw data but on a latent factor compression of those. This indeed bounds the model flexibility and hence the potential for overfitting, and might be more central in this aspect than the ANN model-framework itself.
Cross-validation optimism
When the number of calibration samples is limited, an adequate cross-validation optimism is critical for using RMSECV as a reliable estimate for model validation. In such cases, samples are not required for an independent test set which reduces the number of samples required for model building and validation. However, RMSECV was also used for selecting appropriate meta-parameters so, the RMSECV values can be slightly optimistic. The three calibration methods demonstrated consistent CV optimism profiles across datasets (Figures 1(b), 2(b) and 3(b)), indicating that the effect of sample size on CV optimism follows a certain pattern regardless of the type of data. However, PLS models of the QSAR data (Figure 2(b)) illustrated that exceptions do exist.
SVR obtained equally low CV optimism as PLS, even though RMSECV was used to select three meta-parameters rather than the single parameter for PLS. The negligible CV optimism of the PLS and SVR models implied that the model performance could be evaluated using cross-validation and thus the need for setting aside samples for a test set was avoided. This is especially valuable in practical applications where data are limited due to e.g. cumbersome or expensive sample collection as well as attainment of reference measures.
ANN obtained positive CV optimism in general. This resulted from a more extensive use of the cross-validated error for parameter selection during ANN model building compared to PLS and SVR: selection of meta-parameters as well as deciding the number of learning cycles for ANN parameter estimation was all based on RMSECV. While this extensive use of RMSECV had no implications for large datasets, which held sufficient information to estimate the model parameters accurately, it biased RMSECV downwards for small datasets. However, it had only minor influence on the models, since ANN performance was only a little poorer than the SVR performance for similar calibration data sizes.
As a consequence of the high CV optimism, proper evaluation of ANN model performance required test set validation. This pointed at another disadvantage of ANN: in addition to requirement of a larger calibration sample set in order to obtain model accuracy similar to SVR, it also required samples set aside for evaluating the true model performance.
High dimensional data
Our study revealed a moderate effect of sample size on model performance although two of the three data types were of very high dimensions. The only extreme result was obtained for the linear PLS on the highly complex QSAR dataset and not for the more flexible nonlinear methods, indicating that the combination of high dimensions and high model flexibility, even for small sample sizes, did not radically impair model prediction performance. This may be due to the strategies for data reduction either by latent factor compression (ANN) or application of kernels (SVR), which indeed restrained the curse of dimensionality.
Conclusion
In this paper, we applied PLS, ANN and SVR to three very different datasets containing nonlinearities in order to study the effect of calibration sample size on model prediction performance. The nonlinear methods performed equally well or better than PLS for even small sample sizes and in general only moderate effects of sample size were observed. Hence, our results challenge the dogmatic perception of ANN requiring a high number of calibration samples in order to perform well. SVR obtained consistently low CV optimism and superior model performance which pointed at a high generalization ability, even for small sample sets. Two datasets were of high dimensions, which illustrated that the combination of high dimensions and high model flexibility for even small sample sizes did not radically impair ANN and SVR prediction performance.
We demonstrated the applicability of nonlinear calibration methods for real world applications where the amount of calibration data can be limited. Our results indicated that practitioners have no need to hesitate on using ANN and SVR due to concerns about sample size requirements. Particularly, our results pointed at SVR as a good candidate for real world applications, as it showed a number of advantages over ANN as implemented here: (i) generally it obtained higher accuracies, why fewer calibration samples were required to obtain a desired model performance and (ii) for small calibration datasets, it was less sensitive to sampling variation. However, smaller sample sizes were in general associated with higher variance in RMSEP. Finally (iii) SVR was robust against CV optimism for even relatively small sample sizes, so cross-validation was a realistic option for validation of the true model performance. In a case of limited availability of samples, this would significantly reduce the sample requirement since a test set would not be needed. Additional advantages of SVR are its global solution, its deterministic behaviour and a more straightforward meta-parameter selection although selection of optimal meta-parameters may require substantial computer power and be highly time consuming.
Footnotes
Acknowledgement
The authors acknowledge Daito Togyo Corporation, Japan for providing the process samples.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
