Abstract
Infrared spectroscopy is a prominent molecular technique for bacterial analysis. Within its context, near infrared spectroscopy in particular brings benefits over other vibrational approaches; these advantages include, for example, lower sensitivity to water, high penetration depth and low cost. However, near infrared spectroscopy is not popular within microbiology, because the spectra of organic samples are difficult to interpret. We propose a comparison of spectral curve-fitting methods, namely, techniques that facilitate the interpretation of most peaks, simplify the spectra and improve the prediction of bacterial species from the relevant near infrared spectra. The performances of three common curve-fitting algorithms and the technique based on the differential evolution were compared via a synthesized experimental spectrum. Utilizing the obtained results, the spectra of three different bacterial species were curve fit by optimized algorithm. The proposed algorithm decomposed the spectra to specific absorption peaks, whose parameters were estimated via the differential evolution approach initialized through Levenberg-Marquardt optimization; subsequently, the spectra were classified with conventional procedures and using the parameters of the revealed peaks. On a limited data set, the correct classification rate computed by partial least squares discriminant analysis was 95%. When we employed the peak parameters for the classification, the rate corresponded to 91.7%. According to the Gaussian formula, the parameters comprise the spectral peak position, amplitude and width. The most important peaks for bacterial discrimination were identified by analysis of variance and interpreted as N–H stretching bonds in proteins, cis bonds and CH2 absorption in fatty acids. We examined some aspects of the behaviour of standard curve-fitting algorithms and proposed differential evolution to optimize the fitting process. Based on the correct use of these algorithms, the near infrared spectra of bacteria can be interpreted and the full potential of near infrared spectroscopy in microbiology exploited.
Keywords
Introduction
In microbiology, the accurate and rapid identification of microorganisms constitutes an essential task for diagnosis and treatment. Currently, microbiology laboratories utilize a wide range of techniques to enable classification at the level of the genus and species; these approaches exploit the traditional phenotypic, gene sequencing (polymerase chain reaction, PCR) 1 and proteomics analysis (e.g. matrix-assisted laser desorption/ionization). 2 Although the given procedures are well established, they suffer from certain drawbacks, including considerable time consumption, high costs, excessive complexity and sensitivity to matrix components. With the advent of molecular methods, some of the disadvantages can be overcome by molecular vibration techniques, such as near infrared (NIR) spectroscopy. This method does not require lengthy sample preparation; the cost of the analysis rests only in the filtration membrane (if used), and the matrix influence can be suppressed by the correct interpretation of the NIR spectra or via spectral subtraction techniques. 3
In numerous papers,4–6 infrared (IR) spectroscopy coupled with multivariate statistics is considered a suitable approach to characterizing microbial samples. Basic cellular components, including proteins, fatty acids and polysaccharides, absorb several wavelength bands of IR radiation; the ratios of these components are used to identify a wide range of organisms. The chemical bonds presented in the related substances (C–H, C–N and N–H) are well detectable by NIR spectroscopy, and the sensitivity to such substances allows NIR spectroscopy to be routinely used within the food processing industry and agriculture. Because the ratios of these NIR-active molecular bonds presented in the cells are also species dependent due to a unique cell metabolism, some studies7–9 have reported successful bacterial identification according to NIR spectra; NIR spectroscopy is sensitive mainly to changes in the membrane structure (Gram staining). 10 The variances in the teichoic acid, carbohydrates and amino acids present in the peptidoglycan layer enable us to distinguish between Gram-positive (a higher content of peptidoglycan and a lower content of lipids) and Gram-negative species (a lower content of peptidoglycan and a higher content of lipids, with teichoic acids absent). However, an NIR region comprises broad and overlapping combinations and overtone absorption peaks, which render the NIR spectrum difficult to interpret; because of this factor, then, multivariate statistics is commonly utilized for the classification of bacteria. But the complex structures of biological samples usually do not allow us to process the spectra via basic multivariate statistical models, and therefore some advanced statistics need to be utilized. One of the methods for processing complex spectra is spectral decomposition by means of curve-fitting algorithms.11,12 This tool embodies three major benefits. First, it enables us to interpret the individual peaks, simplifying the differentiation between bacterial and other components (matrix, filter membrane). Second, the classification models are based on dimensionally reduced data (the models classify only the peak parameters, not the whole spectrum), which then improve the model generality and, consequently, the performance. More information on the generality of the classification models is available from the referenced literature.13,14 Third, we can point to the substantial spectral noise reduction, 15 an important precondition for any effective monitoring of weak signals. By contrast, then, the performance of spectral curve fitting depends on many input parameters, whose optimal settings constitute a complicated task; in this paper, we introduce the computational strategy for such settings.
We set up an artificial experimental spectrum with six overlapped peaks. As the positions, widths and amplitudes of each peak were known, we used the spectrum to evaluate different curve-fitting algorithms (the conventional and machine learning-based ones) and their settings. Exploiting the acquired data, we applied the best performing curve-fitting algorithm to the real NIR spectra of bacterial cells. A classification model calibrated with the parameters of the extracted peaks was assembled, allowing us to interpret most of the spectral peaks and to determine the importance of individual peaks for the classification of bacterial species.
The objective of the studies reported in this paper was to find a suitable spectral curve-fitting tool for bacterial cell classification and optimal parameters of the proposed algorithm. For this purpose, the artificial experimental spectrum was created, and the optimized algorithm was analysed using the spectra of real bacterial samples. The interpretation of found peaks is introduced in the second part of the study.
Materials and methods
Microorganisms and inoculum preparation
A collection of three typical food-borne pathogens and food hygiene-related microorganism indicators from the Spanish Type Culture Collection (CECT), including Escherichia coli (CECT 686), Listeria ivanovii (CECT 913) and Salmonella enterica (CECT 699), was examined. The species were stored in the form of colonies on tryptic soy agar at 7℃; the storage time of the agar plates was monitored and maintained in order to control the growth phases of the inoculums. We used a sterile loop to transfer the cells from the centre of a colony to 50 ml of tryptic soy broth (TSB). All the inoculums were left to grow for 24 h, with shaking at 37℃. To suppress the effect of the supernatant, all the bacteria were grown in individual tubes, applying the same batch of TSB (Oxoid, UK) for 24 h with shaking. Prior to the measurements, the bacterial biomass was double washed in a saline solution (0.9%). The rinsing included centrifugation (6000 r/min for 15 min). The subsequent procedures were performed on 10 ml of a suspension exhibiting the same concentration level ensured by levelling to the optical density of 1.5 (measured at the wavelength of 600 nm with a Thermo Fisher Scientific Helios α and a 10-mm cuvette). The concentration of the suspension was determined to be 109 CFU ml-1 via standard plate count, using plate count agar (Oxoid, UK). Such a level ensured sufficient signal for bacterial analysis. Within the research described in this paper, the lowest rate of concentration to facilitate classification was not determined; however, the bottom limit was proposed to equal 108 CFU ml-1 in a related article. 10 The suspension was filtrated using vacuum filtration system (Sartorius, Germany) by standard procedure. 10 As we used 10 ml of the above-defined suspension with 109 CFU ml-1, approximately 1010 cells were located on the filter. Due to the fact that all the three bacterial species were rod shaped and exhibited a size larger than the filter pores, we assumed that the quantity of cells filtered by the membrane reached very similar levels in all of the species. For the NIR analysis, in each bacterial culture, the suspension was filtered with a glass fibre filter exhibiting the pore size of 0.4 µm (Macherey-Nagel, MN GF-5, 47 mm in diameter; ftp://ftp.mn-net.com/english/Flyer_Catalogs/Filtration/Cat_Filtration_notitleEN.pdf) and dehydrated in an incubator at 60℃ for 2 h to eliminate the effect of water. The drying procedure removed all the unbound and a part of the bound water in the cell membrane, 16 and thus it exerted a rather complicated influence on their survival, depending on a multitude of factors (such as the growth stage and species); however, the chemical composition of the cells remained stable. The drying affected mainly the structure of the cellular components that had only a minimal effect on the NIR spectrum. 17 Importantly, since we did not expect the bacterial cells to survive, the NIR measurement was considered destructive. The moisture in the filter was easily detectable according to the peak at 1950 nm, and we thus found the optimal time during the initial experiments (2 h). As most of the NIR measurements were not carried out immediately after the drying process, we kept the filters in a Petri dish containing silica gel and sealed with parafilm (http://www.parafilm.com/); these preliminary measures then enabled us to maintain a relative humidity of below 30%, verified by the colour indicator in the applied silica adsorbent. It is vital to measure the bacterial biomass within the same growth phase because, considering the individual impacts upon the bacterial chemical composition, the time of growth constitutes a factor markedly more significant than the species variance. This phenomenon was discovered during the initial experiments and is discussed in the related paper by Al-Qadiri et al. 18 The procedure was repeated several times during 12 days, accumulating a total of 60 samples for each microorganism.
Real samples are usually not in the form of a monoculture; to isolate a pure bacteria culture from a mixed sample, we can employ the streaking method 19 or inoculation on a suitable selective agar.
NIR analysis
The NIR spectra were acquired via a Nicolet Antaris Fourier transform-NIR analyzer (Thermo Scientific, USA; www.thermofisher.com). Because of the scattering character of the sample, the diffuse reflectance technique with an integrating sphere was used. The advantage of this approach lies in collecting transmitted and scattered light, which is necessary for the measurement of samples that cause light scattering. During the run of the diffuse reflectance technique, an incoming beam was reflected, scattered and transmitted by the sample. As the sizes of the bacterial species are within the micrometer range, significant light scattering was observed, this effect is roughly describable using Mie theory. 20 The scattering added a tilted baseline to the spectra, but this can be mathematically suppressed (as proposed below). The filters were placed on the integrating sphere window and covered with a reflecting aluminum plate. The spectra were recorded within the range of between 9091 and 4000 cm−1 at intervals of 1 cm−1, corresponding to the range of 1100 to 2500 nm measured in 6224 wavelengths (data points) at the varying wavelength interval of between 0.12 and 0.62 nm. The acquired spectrum was averaged from 50 scans. All of the accessories were cleaned with 70% ethanol between the measurements. In total, we measured 180 spectra, with 3 spectra acquired from each filter. Because multiple measurement of one filter is considered technical repetition, where the data independence condition remains unfulfilled, the data set was reduced to 60 spectra via averaging the 3 spectra acquired from one filter. Conversely, the original data set was beneficial for curve-fitting algorithms, and it was intentionally used for the second part of the study.
Spectra treatment and statistical analysis
Before creating the classification and exploratory models (principal component analysis, PCA and partial least squares discriminant analysis, PLS2-DA), we treated the spectra with multiplicative scatter correction 21 or standard normal variate, techniques that suppress the spectral variation caused by light scattering on the bacterial cells. These methods provided only minor improvement in the classification rate (1%) because wavelength-dependent scattering influenced mainly the shift and tilt of the baseline; these effects were, however, sufficiently suppressed by spectral derivation (as proposed below). For this reason, we did not utilize the mathematical suppression of the scattering effect.
The spectral variation arising from the different ratios of the main chemical constituents among the bacterial species was negligible; thus, techniques for spectral resolution improvement, such as the first derivative of the spectra, were applied. Due to the lower absorption of all the samples of bacterial species in the region of 1100–1683 nm, this range was noisy and here we applied the second-order Savitzky-Golay filter
22
with the window size of 111 data points. For the remaining part of the spectra, the window size was determined to be 71 data points. This approach ensured noise elimination across the whole range, with the spectral contours preserved. The raw and the resultant spectra are illustrated in Figure 1 and Figure 4 respectively.
The acquired NIR spectra of the bacterial samples. NIR: near infrared. The second derivative of the spectra and the interpretation of selected bands. The inset shows the partial least squares discriminant analysis (PLS2-DA) weight plot.

Classic multivariate models, such as PCA and PLS2-DA (described by Barker and Rayens 23 ), provided the basic explanatory and classification analysis. For a proper validation of all proposed models, we utilized the two-fold cross validation (CV) scheme. In this procedure, the randomly assigned 12 spectra from each species were considered a validation set, and the model was calibrated according to the remaining set of spectra. In the second iteration, the sets were swapped. The correct classification rates (CCR) were evaluated and compared. All the statistical techniques were implemented using Matlab 8.1a. (The Math Works Inc., Natick, MA, USA)
Spectral curve fitting
The following part of the study presents an algorithm for spectral curve fitting. The curve-fitting approach facilitates the interpretation of the given spectra and can lead to a simplified classification model; moreover, the location of fundamental spectral peaks enables us to directly utilize the Beer-Lambert law.
24
The analytical expression of the spectral domain bands is usually not feasible, since the spectral line broadening effects cause massive peak overlapping and make the direct identification of the peak positions difficult to achieve; thus, numerical solution becomes a necessary step. The calculated spectrum consisted of 31 peaks, including their positions, amplitudes and widths of the Gaussian shape profile. The final count of the peaks was determined by the trial-and-error approach: A lower number of peaks did not approximate the measured spectra correctly (the mean deviation between the curve fit and measured spectra was more than 4%), and any number higher than 31 did not result in further improvement of the curve-fitting precision. The Gaussian peak shape is given by Heisenberg's uncertainty principle in energy,
25
Doppler’s broadening by moving molecules
25
and other broadening effects caused by molecular proximity and collisions. As the molecular vibration coherence lifetime is high (comparable to the molecular relaxation time), the excited molecules relax quickly according to the Gaussian distribution. The measurements were carried out exclusively with solid state samples, eliminating the necessity to perform additional approximation, in which the Voigt or Lorentzian profiles are employed. A detailed description of techniques for the curve fitting of vibration spectra and their physical background is proposed in relevant references.26,27 The Gaussian curve was implemented according to equation (1). Using the parameters a, w0 and σ, we can adjust the shape of the waveform (the amplitude, position and width). The spectral wavelengths are expressed by the symbol w. We have
The modelled spectrum consists of superposed Gaussian curves. The superposition is shown in equation (2), where P is the number of peaks presented in the modelled spectrum. Because the Gaussian curve is given by three variables (the amplitude, position and width, as indicated in equation (1)), we need P*3 independent variables to compose the spectrum. These variables are retrieved via an optimization procedure based on the criterion minimization specified in equation (3), where Tm stands for the original measured spectrum of the bacterial cells. The criterion can be computed using various derivations (α). We then have
In order to identify the correct peak parameters corresponding to the lowest criterion value, we are introducing common curve-fitting procedures, such as the Nelder-Mead simplex method (NMSA), 28 Levenberg-Marquardt method (LMA) 29 and trust-region-reflective least squares algorithm (TRRA). 30 The proposed algorithms are essentially non-linear regression models consisting of optimization techniques that adjust the parameters of the peaks until the best agreement with the acquired spectra is achieved. More detailed description of the techniques is available within the above-referenced papers. In our research, these methods are compared with the robust procedure embodied in the differential evolution (DE) algorithm. 31
Differential evolution algorithm
One of the methods proposed to minimize the criterion is the DE algorithm. DE is a stochastic, population-based optimization algorithm related to the concept of evolutionary algorithm, whose details are outlined within the corresponding referenced papers. 32 , 33 Advantageously, DE can be used for non-differentiable, non-linear problems. As the actual spectral decomposition tends to have more solutions, the optimal option can be chosen from a series of spectra proposed by the algorithm. This feature is especially useful for spectral decomposition because the solution can be selected according to some prior knowledge about the measured spectrum.
DE contains all evolutionary principles, including mutation, selection and recombination. The optimized variables are presented in a matrix Px3. The algorithm operates within the following steps:
Initialize N matrices of parameters - Mutation. Assemble a donor matrix
The new matrix is modified as shown in equation (4); here, F is the input parameter known as the mutation factor, and it needs to be anteriorly defined in the interval 0-1. We have:
Recombination. The compilation of a trial matrix ( Selection. The criterion value of the original matrix The recombination and selection are repeated, while the criterion value of the best solution decreases; various terminating conditions can be implemented (for example, time constraint or the lowest achieved improvement in the criterion value; more information is available from the referenced literature
34
).
Results and discussion
NIR spectra
At the first stage, we employed the diffuse reflectance technique to analyse the spectra of the three bacterial species measured in the form of dehydrated suspensions on the filtering membranes; the suspensions were filtered 24 h after inoculation. The acquired spectra exhibited the characteristic bands of organic samples, as shown in Figure 1; the bands were roughly interpreted in numerous sources.35,14 The main bands can be defined as follows:
Carbohydrates: 1483 nm, 1490 nm, 2100 nm; Proteins: 2050–2060 nm, 1500–1530 nm; Fats: 2222 nm, 2070 nm, 1203 nm; Moisture: 1947 nm, 1450 nm.
The peaks were wide and overlapped, and the borders between them remained difficult to find. The readability of the spectra was also decreased by the heavy baseline tilt. However, by further extension, the identification of concrete spectral contours in bacterial samples was analysed within several related papers7,9,10 and will be discussed below.
The optimal growth time was determined to be 24 h, mainly for practical reasons and respecting the growth curves of all the species investigated; this period corresponded to the stationary phase of bacterial growth. Although some research reports 18 indicated the possibility of comparably successful bacterial identification in the later phases, it is essential to avoid merging the different growth stages into one classification mode. At this point, let us note that the NIR spectrum of the bacterial biomass kept changing during the different growth stages; a detailed investigation of this phenomenon is outlined in a study of Al-Qadiri et al. 18
Principal component analysis
The output of the exploratory analysis shows the PCA score plot (the first two components) presented in the upper section of Figure 2; the clearly distinguishable cluster relates to the examined Listeria ivanovii. Unlike the other species involved in our research, Listeria ivanovii is a Gram-positive bacterium with a different composition of the cell wall (a thicker peptidoglycan layer); hence, we expected a divergent spectrum. The contributions of all PCA components are expressed by eigenvalues of the PCA correlation matrix. The eigenvalues are shown in Table 1.
The principal component analysis (PCA; top) and partial least squares discriminant analysis (PLS2-DA; bottom) score plots of the preprocessed NIR bacterial spectra. NIR: near infrared.
Classification
The eigenvalues of the PCA correlation matrix.
PCA: principal component analysis.

The CCR and cumulative variance of the partial least squares discriminant analysis (PLS2-DA) and principal component analysis (PCA) methods. CCR: correct classification rates.
The data set preparation and performance of the PLS2-DA model.
CCR: correct classification rates; CV: cross validation; der.: derivative; MSC: multiplicative scatter correction; PLS2-DA: partial least squares discriminant analysis; SG: Savitzky-Golay filter.
The spectrum is partially interpretable according to the PLS2-DA loading (refer to the inset of Figure 4). Some basic detectable cellular components are highlighted in Figure 4, and high values of the PLS2-DA loading are present mainly in the range of between 2000 and 2500 nm. The N-acetylmuramic acid from the peptidoglycan layer leads to high values of the weight vector of the first component within the band limited by 2200 and 2500 nm; this corresponds to our assumptions because the first component separates Gram-positive bacteria from Gram-negative ones (Figure 2).
Even though only three different species were distinguished, it can be seen that the NIR spectrum contains information about the diverse chemical compositions of the bacterial cells. Within the appropriate phase, we utilized the curve-fitting decomposition to simplify and fully interpret the bacterial spectra.
Curve fitting of an experimental synthetized spectrum
To evaluate the curve-fitting techniques, an experimental spectrum comprising six peaks was established in order to exhibit a degree of overlapping similar to that of the original spectrum (Figure 5). Because the baseline tilt of the original spectra was eliminated by applying the second derivative, it was not modelled in the experimental spectrum. As the parameters of the peaks were known, it was possible to assess the precision of all the curve-fitting functions. We investigated the potential of the conventional and DE algorithms under various initial values, criterion formulations and noise conditions. A simple optimization criterion can be calculated according to equation (3), or more comprehensively, based on equation (6); we then have
The experimental spectrum to evaluate the curve-fitting methods and their parameters. The solid line represents six created peaks and the dashed line represents their superposition.
where
N is the number of the spectral derivatives;
In the equation, the criterion is calculated as the sum of the differences between the original and the fitted spectra in N derivatives.
Impact of initialization on the curve-fitting precision in conventional algorithms
The average distances between the fitted and the real peaks computed from the spectra composed via the three curve-fitting methods applied.
Note. The values depend on the initialization method and degrees of derivation present in the curve-fitting criterion.
Since the resultant spectra depend on the initialization, four different initialization sets were tested. The initialization based on the minimum of the second derivative constitutes the standard, most common technique. Additionally, we propose initializing via equidistant distribution and the location of peaks near and far from the real peak positions. Table 3 shows how the actual initialization affects the precision of the fitting procedure. The positions of the revealed peaks were not optimal, allowing for significant deviations from the real peak positions. The initialization according to the minimum of the second derivative exhibited considerably lower deviations. Equidistant initialization also yielded reasonable results, but some peaks (1614 and 1639 nm) were not precisely estimated. Interestingly, because of its spectral dimness, the peak at 1570 nm was not detected by any initialization set-up. All the procedures identified a false peak around 1590 nm. DE is not included in this comparative study, because it was initialized randomly. For better revelation of the peaks, we considered a different form of the minimized criterion that highlights spectral peaks, namely, derivative selection. The influence of derivative selection on the precision of spectral curve fitting is discussed later in this chapter.
One of the widely applied optimization procedures is the NMSA. However, this general technique appears to be unsuitable for spectral curve fitting; the algorithm gets stuck in local minima, where it performs a multitude of iterations with negligible improvement of the function value. The peaks were not identified correctly even at the expense of very long computational time. In contrast to the other examined techniques, the NMSA is a heuristic method without any convergence theory to solve non-linear least squares problems. We did not find any reason why to prioritize the NMSA over the other proposed spectral curve-fitting methods; thus, the method is not discussed below.
Relationship between the precision of curve-fitting algorithms and derivation selection
Table 3 compares the results produced by various degrees of derivation. The lowest deviation value exhibited curve fitting with the criterion including the first to third spectral derivatives. The results acquired with DE are comparable to those obtainable via conventional algorithms, yet only at the cost of considerably longer computational time. More concretely, in the given case, the computational time rose from 21 s (LMA) to 31 min (DE) on a computer with Intel Core i7, 2,3 GHz. There are two main arguments in favour of utilizing DE for NIR spectra curve fitting. First, and perhaps the most importantly, DE offers a set of various results from which the optimal solution can be chosen. As is obvious from the experimental spectrum, the peak at 1570 nm was difficult to identify due to being overlapped by the much stronger near peak at 1614 nm. If prior information about this peak is available, the solution that includes the peak can be selected.
Relationship between the precision of curve-fitting algorithms and the number of modelled peaks
The averaged distances between the fitted and the real peaks computed from the spectra composed via the three curve-fitting methods applied.
Note. The values depend on the number of modelled peaks. Lower value means better fitting of modelled and real peaks. LMA: Levenberg-Marquardt method; NMSA: Nelder-Mead simplex analysis; DE: differential evolution algorithm.
Relationship between the precision of curve-fitting algorithms and the level of spectral noise
Due to the effect of the NIR instruments, the corresponding spectra are noisy. More information about the noise in NIR spectrometry is available from chapter 3 in the relevant references source.
15
Thus, higher derivatives must be utilized with care because they tend to emphasize the spectral noise. The experimental spectrum was spoiled by noise at different levels, and we investigated the algorithm behaviour. Figure 6 shows a comparison of the LMA and DE in various derivations. Although the LMA was found to be less sensitive to noise, both the algorithms exhibited certain deterioration of prediction when the criterion considered only higher derivatives (the second–third ones). The criterion including lower derivatives (the zeroth–third ones) appeared better in handling noisy data, but all the curve fittings started to be unstable when the signal-to-noise ratio dropped below 25 dB.
A comparison of the Levenberg-Marquardt method (LMA) and differential evolution algorithm (DE) approaches to the curve fitting of a signal with different signal-to-noise ratios.
Curve fitting of the NIR spectra of bacteria
Based on the earlier findings, we fitted the previously measured NIR spectra of bacteria. The spectra were divided into four sections (Figure 7). In each partition, the curve-fitting procedure was applied separately; this approach then led to an approximation precision better than that achieved with the simultaneous processing of the entire spectra. As apparent from the comparison of the curve-fitting algorithms outlined in previous parts of the paper, the NMSA method is not suitable for the fitting of complex waveforms (Table 3). Due to the similar results produced by the TRRA and LMA methods, we introduce a comparison of the LMA and DE algorithms in the next subsection of the paper. The optimization criterion contained the second and third derivatives. It is reasonable to compromise between the immunity to noise and the ability to model featureless peaks hidden through overlapping (see the above section titled ‘Relationship between the precision of curve-fitting algorithms and the level of spectral noise’). Despite the fact that LMA performed worse with equidistant initialization than with that exploiting the minimum of the second derivative (Table 3), we applied equidistant strategy because the other technique proved inefficacy due to the complexity of the spectrum. By further definition, when the spectrum consisted of many peaks located close together, the second derivative minima did not express the original peak positions.
The curve fitting of the second derivative spectrum of E. coli via differential evolution algorithm (DE; solid line – the original spectrum; dotted line – the curve-fit spectrum). The precision of the curve fitting of the other species is at the same level; the spectra vary mainly in the peak amplitudes and, to a lesser extent, the location.
Although the LMA having the above-mentioned parameters ensured good approximation, the spectrum computed with DE better fit the original bacterial spectrum.
The precision rates formulated as the mean deviation between the curve-fitted and the original spectral points are as follows:
LMA: 3.25% (speed: 21 s with an Intel i7 processor); DE: 1.57% (speed: 31 min with an Intel i7 processor).
Note: As the precision rates were similar in all the measured bacterial species, we introduce the averaged precision values.
Because LMA remains incomparably faster, we propose a hybrid approach where the LMA serves for the initialization of DE only; this solution combines the rapidity of the LMA and the advantages of DE. With this combined technique, we reached the averaged precision rate of 1.71% for all the 180 measured spectra (speed: 124 s on an Intel i7 processor). By fusing the favourable behaviour of DE, discussed above in Material and methods, and the rapid approximation inherent with LMA as the initializing algorithm, we reduced the computational time from 31 min (DE) to 124 s (DE + LMA); the actual precision was thus preserved. Special care must be devoted to keeping the DE population diverse; otherwise, some additional algorithm to maintain the diversity needs to be implemented. 36 In this study, we controlled the population heterogeneity by modifying algorithm parameters such as the crossover ratio and the mutation factor (refer to equations (4 and 5)).
The peaks with the lowest one-way analysis of variance P values.
Another distinctive peak located at 2310 nm was associated with the CH2 structure in fatty acids. According to this peak, E. coli could be distinguished (P = 10−7). The peaks associated with the fatty acid were found to be helpful for bacterial discrimination also in other studies. 8
This part of the study presents a list of observed peaks in the spectra of all bacterial strains. One of the most distinctive peaks was located at 1927 nm, and it associates with the O–H combination of water. Due to the high absorption, this peak was distorted by noise. The absorption signal related to the C = O bond in amide was identified at 1982 nm. 7 The band around 2053 nm is connected with the important N–H symmetric and asymmetric stretching modes in amides present in proteins.9,10,39 The weaker peak at 2164 nm is related to bonds in amide.10,39 This region is followed by stronger absorption bands observed at 2184 nm (a cis bond in unsaturated fatty acids 8 ) and 2267 nm, which could correspond to the C–H combination of polysaccharides. 10 The N-acetylmuramic acid from the peptidoglycan layer gives rise to the band at 2299 nm. 10 The CH2 bending and stretching combination bands match the frequency of 2310 nm.8,10,39 Many peaks were noticed in the spectral region of between 2330 and 2368 nm, the molecular vibrations responsible for absorption in this range were associated with C–H aliphatic groups. 9 The bands at 1700–1759 nm are related to the CH3 and CH2 stretching first overtone.10,39 The strong region at 1660 nm is attributed to the C–H stretching first overtone.7,8,10,39 The N–H first overtone dominates at 1463 nm.8,39 Most of the proposed peak interpretations are based on the measurement of separate cellular constituents. All the listed peaks are highlighted in Figure 1, and their locations, represented by grey lines, can be found also in Figure 7.
Further, the computed peak parameters (the location, amplitude and width) were classified according to the bacterial species. For the classification, we used PLS2-DA, and the verification was performed utilizing the two-fold CV. On the introduced data set, we evaluated the CCR at 91.7%. Good classification performance was maintained despite the simplification of the spectra. Such simplification was basically reached by substituting the 6224 spectral points with the 93 peak parameters. The importance of the computed peak parameters for the spectral classification was described by the ANOVA P values. Some peaks having high P values exhibited a variability not caused by the various bacterial species. The parameters with the lowest P values are listed in Table 5.
Conclusion
Curve fitting facilitates the understanding and interpretation of the NIR spectra of bacteria. The behaviour of all curve-fitting algorithms depends on multiple factors, including the noise, initialization procedure, criterion form, and number of modelled peaks. These factors must be considered during the selection of the curve-fitting method. The DE algorithm proved suitable for this task because it exhibits high tolerance to such factors. In particular, the combination of DE and the LMA (the latter technique is used to perform the actual initialization) seems to be beneficial in terms of the rapidity and accuracy of the algorithm. Once the spectrum is decomposed to Gaussian peaks, most of the peaks can be assigned to the main cellular components (proteins, fatty acids and cell wall elements). Utilizing the proposed procedure, it is possible to finally partially interpret the NIR spectra of biological samples and to quantify the benefit of spectral regions for discrimination. According to the computed PLS2-DA weight vectors and curve-fit peaks, the most important spectral regions for bacterial discrimination are in the range of 2000–2500 nm, where the bonds present in protein and fatty acids cause strong peaks; such aspects include N–H stretching (2220 nm), the cis bond in unsaturated fatty acids (2184 nm) and the CH2 stretching bonds in fatty acids and proteins (2310 nm). While the original spectra were classified according to the bacterial species, with the CCR of 95%, the classification based on the parameters (the location, amplitude and width) of the modelled Gaussian peaks reached 91.7%. This result confirms that simplifying the spectra has only a minor impact on the classification performance. The high classification rates can be obtained only using non-mixed bacterial cultures, with a concentration of more than 108 CFU ml-1. The measurements must be carried out at the same bacterial growth stage, and this condition appears to be a crucial part of the sample preparation process. We believe that even though the analysis of merely three species does not fully prove the usability of NIR spectroscopy for bacterial identification, it demonstrates the potential of curve-fitting methods in the simplification and interpretation of the NIR spectra of bacterial cells.
The future research will be focused on the impact of bacterial growth phases on the performance of the standard classification models and curve-fitting procedures. In this context, we also intend to improve the algorithms for DE-based spectral curve fitting and to propose tools facilitating better interpretation of bacterial cell spectra via conventional spectral analysis methods.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The actual research was partially supported by grant No. LO1401 covered by the National Sustainability Programme. The infrastructure of the SIX Center was utilized to facilitate both the research and the assistance provided by the general student development project being executed at Brno University of Technology and guaranteed by the Ministry of Agriculture of the Czech Republic (project Enteropig - QJ1310258). The results and outcomes were prepared using instrumentation funded from the Research and Development for Innovations Operational Programme, project CZ.1.05/4.1.00/04.0135 Teaching and Research Facilities for Biotechnological Disciplines and Extension of Infrastructure.
