Abstract
Maca (Lepidium meyenii Walp.) is a cruciferous edible and medicinal plant rich in nutrients. As maca demand in the international market is gradually increasing, dishonest people have been using low-priced alternatives to either adulterate or falsify maca and increase their profit. Existing methods to identify and quantify adulterated maca are laborious, expensive, destructive, time-consuming, and environmentally unfriendly. Thus, it is imperative to develop a method to overcome these problems to effectively authenticate maca products. We combine near infrared spectroscopy with chemometrics to classify and quantify maca powder adulteration by turnip and radish powder. Different maca samples were adulterated with turnip and radish powder individually at different percentages (5–95%). Specifically, discriminant analysis based on a support vector machine provides a classification accuracy of 100%, allowing near infrared spectroscopy to be used to distinguish maca powder adulteration. Furthermore, to calibrate a regression model, we evaluated the partial least squares (PLS), interval PLS, and synergy interval PLS (siPLS). The siPLS models were determined as the best models for the quantification of maca powder adulterated with turnip and radish powder, in which the correlation coefficients were both 0.97 with root-mean-square error of prediction values of 5.79% and 5.85% for the two models, respectively. The combination of near infrared spectroscopy and chemometrics can provide a fast, simple, and environment-friendly analytical method for identifying and quantifying the properties of maca.
Keywords
Introduction
Maca (Lepidium meyenii Walp.) is a cruciferous edible and medicinal plant rich in nutrients and native to the Peruvian Andes, which reaches altitudes from 4000 to 4500 m. 1 It contains amide, 2 dietary fiber, 3 glucosinolates and polyphenols 4 and other functional ingredients, which indicate the value and quality of maca. Moreover, studies have shown that consuming the maca root can improve memory, immunity, sperm motility, reduce fatigue, provide anti-oxidation effects, among other benefits.5–7 As maca demand in the international market is gradually increasing, dishonest people have been using low-priced alternatives to either adulterate or falsify maca and increase their profit. In particular, turnip and radish are very similar to maca in terms of appearance and odor and are often incorporated into maca powder for sale. Therefore, ensuring the quality and purity of maca products has become an important problem for the development of its industry.
Current analytical methods to verify the authenticity of maca include gas chromatography–mass spectrometry, 8 Fourier-transform infrared spectroscopy, 9 and DNA barcoding. 3 However, these methods are complex, expensive, and destructive. Although current techniques do not require a complicated sample separation process, it is necessary to extract the functional component (i.e., alkaloid) of maca to then perform rapid detection of maca adulteration using infrared spectroscopy. In general, the roots of maca were extracted with aqueous ethanol, and then the residue was repeatedly purified by column chromatography on silica gel and preparative or semi-preparative HPLC to get the alkaloids. 10 In contrast, near infrared (NIR) spectroscopy provides low cost, fast, and non-destructive analysis. 11 , 12 NIR spectroscopy also avoids time-consuming and labor-intensive sample preprocessing, and detection does not require hazardous chemicals, making it suitable for untrained personnel to perform real-time onsite analysis. NIR spectroscopy registers spectral bands mainly corresponding to C-H, O-H, and N-H vibrations, which are overtone and combination bands with varying NIR frequency regions. 13 By comparing the spectra and analytical composition or nature, a calibration model could be established to quantitatively measure a component or its nature from samples. Through chemometrics, 14 NIR spectroscopy can efficiently extract useful information from a large amount of data and establish a model after obtaining sample measurements. In recent years, combining NIR spectroscopy and chemometrics has become a development trend and is being widely used given its suitable outcomes. In fact, with the development of NIR spectroscopy and chemometrics, improved analyses of agricultural, food, and pharmaceutical products have been proposed.15–18 For example, Ma et al. 19 utilized NIR spectroscopy to determine adulteration of Shan Yao powder, obtaining a correlation coefficient above 0.99. Nie et al. 20 used it for detecting Panax notoginseng powder and its dopants, also obtaining high accuracy. However, to the best of our knowledge, no research has been reported on the use of NIR spectroscopy to detect and quantify maca adulteration.
NIR spectroscopy allows one to obtain large amounts of spectral data in short time. However, the raw spectra acquired from the NIR spectrometer contained background information and noise in addition to sample information. Furthermore, the sample information in some regions is weak and not correlated to the sample composition and properties. Including these data during modeling could cause calculation burden, increased model complexity, and low accuracy. The optimization of independent variables can simplify models and more importantly, by eliminating irrelevant variables and noise, a correction model with strong predictive ability and high robustness can be obtained. Several methods have been considered to remove irrelevant wavelength variables and improve prediction performance. For example, partial least squares (PLS) regression based on latent variable design can satisfy these requirements,21,22 and interval PLS (iPLS) as well as synergy interval PLS (siPLS) are extensions that have been widely used for modeling and prediction. 19 , 23 The model selects the spectral region to enhance correlation with the target sample content. On the other hand, although measurements are important, the employed classification scheme is usually more meaningful, especially for detection and estimation. Support vector machines (SVM), originally proposed by Vapnik, 24 are a widely used pattern recognition technique 25 based on statistical learning theory and structural risk minimization. The SVM algorithm maps data onto a high-dimensional feature space through a nonlinear function (nuclear function), linearizes the nonlinear problem and performs linear classification or regression on this feature space.
In the case of maca, limited effort has been devoted to the investigation of fast techniques to determine its adulteration. Therefore, we propose the identification and quantification of adulterated maca powder with turnip and radish powder by using NIR spectroscopy and chemometrics. Specifically, we established classification models using SVM to distinguish adulterated and pure maca samples. In addition, we developed a calibration model using PLS regression to quantify the adulterant levels and investigated the optimal wavelength for the calibration model by examining iPLS and siPLS besides the conventional PLS.
Materials and methods
Sample preparation
Maca and turnip samples were collected from Yunnan Province, China, whereas radish samples came from a market of Guangdong Province, China. The dirt on the surface of all samples was rinsed, and then they were dried in an oven at 50 °C until reaching a constant weight. We prepared samples of adulterated maca by drying maca, turnip, and radish powder, then passing each powder through a mesh sieve no. 100. The adulterated powders of maca were prepared by mixing the above turnip and radish powders to maca samples at ratios of 5%, 10%, 15%, 18%, 20%, 23%, 25%, 30%, 38%, 40%, 45%, 48%, 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90% and 95% by weight.
NIR spectra acquisition and preprocessing
The samples were scanned using a XDS™ Rapid Content grating spectrometer (FOSS NIRSystems, Inc., Laurel, MD, USA) to obtain their spectra. The spectrometer was equipped with a Si detector (400–1100 nm), a PbS detector (1100–2500 nm), and a diffuse reflection accessory. The spectrum of each sample was collected at 2 nm wavelength intervals from 400 to 2500 nm, thus including the NIR region and most of the visible spectrum. Samples were packed in a quartz vial, compacted and evenly placed in a sample holder for measurements. The spectrum average from three measurements per sample was used for subsequent analysis. Spectra were stored as reflectance (R) data as the logarithm of its reciprocal. During measurements, the laboratory temperature remained at 23 ± 1 °C and 46% of relative humidity.
As the data acquired from the spectrometer contain noise polluting the sample information and possibly interfering with the reliability and stability of a derived estimation model, the raw spectrum of each sample requires preprocessing. Spectral preprocessing methods included standard normal variable 26 and multiplicative scatter correction (MSC), 27 and evaluated the first and second derivative methods. Finally, the second-order MSC was applied on the sample data for subsequent classification and quantitative estimation of the adulteration in maca powder.
Chemometric analysis
SVM classification
SVM is a pattern recognition algorithm based on structural risk minimization
28
providing high classification performance on small sample datasets, and has been used for solving several classification problems.29,30 SVM relies on kernel functions such as the sigmoid, polynomial, and radial basis functions, which are the most widely used.
31
The radial basis function was selected in this study, as it can both map nonlinear data into a higher-dimensional space and reduce the parameters to be estimated. For parameter optimization of this kernel function, we used grid search traversing the combination of penalty parameter c and kernel function parameter g within a specific range to determine the SVM classification model. The optimal parameters correspond to those retrieving the highest classification accuracy from the training set, which contains measurements from 141 samples. Specifically, 22 samples of pure maca, 58 of adulterated maca with turnip powder, and 61 of adulterated maca with radish powder were selected to construct the training set. In addition, 75 samples (12 of pure maca, 29 of adulterated maca with turnip powder, and 34 of adulterated with radish powder) to construct the test set, as shown in Table 1. Furthermore, the NIR spectral parameters assigned to the independent variable matrix
The numbers of samples numbers in training and test sets.
Adulteration quantification using PLS
After identifying adulterated samples using SVM, we applied PLS regression to estimate the proportion of turnip and radish powder to maca powder. PLS 32 provides a strong anti-interference ability and can be used to establish a multivariate calibration model at full wavelength. The predictive ability of the models was evaluated by the correlation coefficient (r), root-mean-square error (RMSE), and the residual predictive deviation (RPD).33,34 RPD, which is defined as the ratio of population standard deviation to standard error of cross-validation for NIR prediction, represents the ability of the model to predict unknown samples, taking into account the variability of the training dataset. Williams and Norris 35 put forward the RPD threshold value in order to understand the calibration possibility for new sample sets by the established quantitative model: if the RPD is below 2.3, the model cannot be used, and if it is above 8, it indicates superior model performance. We evaluated the r in the training and test set, the root-mean-square error of cross-validation (RMSECV), and the predicted root-mean-square error (RMSEP)33,34 of PLS-based models. To obtain the root-mean-square error of cross-validation (RMSECV), we applied cross-validation by removing the spectrum of one sample from the training set, then determining the calibration model using the spectra of the remaining samples in this set, and finally using the removed sample to verify the model. This process was repeated for every sample in the training set.
Wavelength selection
NIR technology can analyze samples in short time by relying on a suitable analytical model. PLS improvements, namely, iPLS 36 and siPLS can correct the relevant characteristic wavelengths of the model, thereby improving the estimation accuracy. Specifically, iPLS was used to divide the wavelength range into 10, 20, and 30 equidistant and nonoverlapping spectral intervals and establish PLS calibration models from each spectral subinterval. Then, we selected the subinterval retrieving the smallest RMSECV. Similarly, siPLS was used to divide the full spectrum (400–2500 nm) into 10, 20, and 30 equidistant spectral subintervals, which were combined in groups of two, three, and four subintervals to establish the PLS calibration model. Likewise, we selected the combination retrieving the smallest RMSECV.
Algorithm implementation
Classification and estimation algorithms were implemented in MATLAB R2014a (Mathworks, Inc., Natick, MA, USA) running on a computer with the Windows 7 operating system (Microsoft Corp., Redmond, WA, USA). For the PLS, iPLS, and siPLS, we used the PLS software package provided by Norgaard, iToolbox. A full cross validation method (Leave-One-Out validation) was carried out to select the optimal number of PLS factors. NIR spectral preprocessing and SVM classification were implemented in MATLAB.
Results and discussion
Spectral analysis
The samples measured by NIR spectroscopy had a large adulterated content range (5–95, mean: 34.3, and median: 25%). Descriptive statistics for all samples indicated that the adulterated content frequency distribution was relatively flat (kurtosis = 2.161) and right skewed (skewness = 0.654). The training and test datasets also yielded similar descriptive statistics.
The raw NIR spectra of pure and adulterated maca samples are shown in Figure 1(a), where the spectral curves of three types of samples are very similar. In fact, the difference between them is extremely subtle to be discerned by the naked eye. Still, spectral preprocessing can unveil the differences caused by adulteration on the samples. Figure 1(b) shows the result of the second-order derivative combined with MSC applied to the pure and adulterated samples, from which it is still difficult to distinguish the spectral differences between them. These results highlight the importance of using chemometric methods for analysis.

(a) Raw and (b) second derivative and MSC preprocessed NIR spectra of pure maca powder (blue) and the same maca powder adulterated with radish (black) and turnip (red).
Discriminant analysis using SVM
Discrimination between authentic and adulterated maca powder was investigated using SVM. The SVM classification model radial basis function was optimized by determining the penalty parameter c and kernel function parameter g. Figure 2 illustrates the grid search for these parameters, from which optimal values for c = 0.25 and g = 588.13 were obtained. Using these parameters, SVM classification achieved 100% identification accuracy of pure and adulterated maca powder when applied to both the training and test sets. Table 2 details the classification performance of the SVM model, where perfect classification and no overfitting were obtained, ensuring the robustness and reliability of the model. Compared to other classification algorithms, SVM does not require additional preprocessing to reduce the dimension of the input space and retrieves a small number of parameters for optimization. Hence, the pure and adulterated samples can be easily separated.

Cross-validated accuracy of SVM model applied to training set according to penalty parameter c and kernel function parameter g.
Classification results obtained from SVM classification.
The proposed SVM classification model in this paper can accurately distinguish adulterated from pure samples at a low cost. Moreover, each sample can be processed in few seconds, improving the applicability for sample monitoring in real settings.
Adulteration quantification
To determine the quality of the variable selection algorithm, we estimated the adulterated powder content using full-spectrum PLS. For the maca adulteration with turnip and radish powder, the PLS model retrieved 6 and 8 latent variables, respectively, as detailed in Table 3.
Statistical results of suitable calibration models to estimate amount of adulterated maca and their PLS models.
LV: number of latent variables.
Note: Significance for the bold emphasis which show that the number of subintervals and the number of synergic subintervals both can influence siPLS modeling effect. Among them, for the turnip adulteration quantification of maca powder, the siPLS model with 20 subintervals and synergic subinterval with three subintervals [4], [8], [18] had the best prediction effect; for the radish adulteration quantification of maca powder, the siPLS model with 10 subintervals and synergic subinterval with two subintervals [4], [7] had the best prediction effect.
The development of spectral interval selection was first conducted on iPLS, whose results are listed in Table 3. The RMSEP values of the full-spectrum PLS model are lower than those of the iPLS algorithm. Moreover, the iPLS calibration models of the turnip- and radish-adulterated maca show the same results, possibly because the information of maca, turnip, and radish is distributed throughout the spectral range, but iPLS only examines individual intervals. Without further examination of the combination of spectral subintervals, there is not enough information to estimate the adulteration ratio in maca, making it difficult to achieve better results than those obtained from full-spectrum PLS regression.
The full-spectrum PLS regression model reflects unrelated variables, which may have a negative impact on the calibration model. In contrast, iPLS only selects a subinterval to develop the model, possibly discarding useful information. Therefore, the judicious selection of the spectral region could improve the predictive power of the PLS model. The siPLS extends the iPLS by combining the subintervals of different local models showing higher accuracy to establish an improved prediction model. The siPLS combined two, three, and four subintervals, whose results are also listed in Table 3. It can be seen from Table 3 that the number of subintervals and the number of synergic subintervals both can influence siPLS modeling effect. Among them, for turnip adulteration quantification of maca powder, the siPLS model with 20 subintervals and synergic subinterval with three subintervals [4], [8], [18] had the best prediction performance. The r and RMSEP values for the prediction model were 0.98 and 5.79 respectively. Likewise, after dividing the spectrum of the radish-adulterated maca into 10 subintervals, the siPLS model with 10 subintervals and synergic subinterval with two subintervals [4], [7] had the best prediction performance. The r and RMSEP values for the prediction model were 0.97 and 5.85 respectively. The scatter plot of Figure 3 shows the correlation between the real and estimated adulteration levels in samples obtained from the best model of the test set mentioned above. From the results, the slopes of both models are close to 1, and the RPD values are above 4, indicating the good predictive performance of the proposed models. Compared with full-spectrum PLS, siPLS improves the prediction ability of the model regardless of the training and test sets, and the amount of both modeled spectral variables and latent variables is notably reduced.

Measured against predicted levels (%) of turnip and radish content in maca powder using NIR spectroscopy and siPLS. (a) Turnip and (b) radish test sets.
The bandwidth intervals selected for turnip-adulterated maca belong to three regions, namely, 746–854, 1168–1270, and 2194–2294 nm corresponding to the third overtone of C-H and hydroxyl group, the second overtone of C-H, and the combination of C-H, respectively. The bandwidth regions for the radish-adulterated maca are 1058–1262 and 1676–1880 nm corresponding to the second overtone of C-H along with the combination related to amylose and the first overtone of the C-H stretch attributed to aliphatic compounds, respectively. Note that the 1850–1880 nm bandwidth corresponds to CONH, which might be attributed to amide compounds.37,38 The bands ranging from 1068 nm–1200 nm were related to the NH stretching modes in most amino acids. The 2398 nm bandwidth corresponds to aromatic compounds, which might be attributed to imidazole alkaloids of maca.
Table 3 shows that the optimal adulteration estimation is obtained using the siPLS method, which also provides the RPD values higher than that of the other two models. This can be explained as follows. In the establishment of the PLS model, its 1050 variables over the spectrum are used to calibrate the model. Among these variables, many independent variables are noninformative and undermine the predictive ability of the model. In contrast, the iPLS method reduces the variables involved in the calibration model. However, selecting only one interval for the iPLS model excludes some valid variables. This problem is solved by using the siPLS method, resulting in better prediction ability than both the full-spectrum PLS and iPLS methods and a simpler model than the full-spectrum PLS method, as confirmed by the reduced number of variables included in the siPLS model compared to the PLS model.
Conclusions
The combination of NIR spectroscopy with chemometrics was able to identify pure and adulterated maca, and establish a fast estimation of turnip- and radish-adulterated maca samples. To identify adulterated maca, SVM preprocessing with grid-search optimized penalty parameter c and kernel function parameter g was used. Evaluation results retrieved a 100% classification accuracy of pure and adulterated samples. After identifying the turnip- and radish-adulterated maca, the adulteration proportion was estimated using NIR spectroscopy with PLS regression. Comparing PLS, iPLS, and siPLS, the calibration model using siPLS provides the best results by notably reducing the complexity of the model and achieving the highest accuracy in both training and test sets. The proposed approach provides a fast, simple, and environment-friendly analysis method to determine the adulteration of maca powder. This approach can be beneficial to foster maca quality control and processing research.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
