Abstract
The challenge of deriving quantitative information from the infrared spectra of proteins arises from the large number of secondary structures and amino acid side-chain functional groups that all contribute to the spectral intensity, such as within the amide I band (1600–1700 cm–1). The band is invariably heavily convoluted from overlapping spectral features, thereby making interpretation difficult such that deconvolution is usually required. This work critically examines the methods available to deconvolute the spectra and assesses the commonly used methods and algorithms applied to vibrational spectra for smoothing and peak identification. We show that unless their spectra have very high signal-to-noise ratios, quantitative analysis to decipher protein constituents is not feasible. The advantages and disadvantages of spectral smoothing using adjacent averaging, the Savitzky–Golay filter and the fast Fourier transform filter are examined in detail. The use of derivative spectra to identify peaks is described with particular reference to the influence and reduction of interfering water bands in the amide I region. The reliability of band narrowing techniques such as second-derivative analysis or Fourier deconvolution that lead to the identification of the contributing protein peaks is investigated. Both methods are shown to be limited in their capacity to resolve features with very similar frequencies. Additionally, the presence of narrow bands arising from high-frequency noise whether from atmospheric water vapor, acoustic vibrations, or electrical interference results in both methods becoming increasingly unusable as narrow bands are preferentially enhanced at the expense of broad ones such as the amide I bands. An optimal strategy is critically developed to allow accurate determination and quantification of protein constituents and their conformations. Additionally, quantitative methods are proposed to account for baseline shifts, which would otherwise introduce significant errors in similarity indices.
Keywords
Introduction
The determination of protein secondary structures or identifying their constituent amino acids is important across a broad range of topics, for example: protein misfolding that is strongly linked with neurodegenerative disorders, tumor cells, red blood cells in sickle cell disease, etc.1–10 The vibrational amide I band in the infrared has proved to be of great significance in the spectroscopy of biological species due to its abundance in constituent proteins and hence amino acid residues.11–13 Confidence in assigning the peak positions that contribute to the overall amide I band as well as accurately fitting the spectrum is of great significance since it allows a more robust determination of the secondary structures that are present as well as their associated quantitative analysis. However, due to the great array of different amino acid residue side chains, the structures of the proteins and the conformations that they adopt under different environmental conditions, the variance in the amide I stretching vibration from proteins, is large. Hence, the spectrum often appears as a broad band with contributions from all of the different environments present.11–15 By studying proteins in specific, well-defined structures, i.e., the α-helices, β-pleated sheets, β-turns, random coils, etc., it has been shown that each of the specific structural environments has a considerably smaller bandwidth associated with their amide I stretching modes.11,15,16 This has allowed the amide I band, as a whole, to be deconvoluted into its contributing bands using peak-fitting procedures and therefore has been used ubiquitously to analyze different protein secondary structures and their relative proportions.17–23 A recurring problem, however, is the specific methodology used to analyze the band and the associated decision of how many components to fit.22–31 Fitting algorithms will achieve the best fit possible for however many peaks they are instructed to assign. A serious limitation exists in that it is easy to ignore contributions that should be included or conversely include fictitious features.
Additional considerations required for spectral analysis of the amide I band include the contributions of water bands (particularly from water vapor) and the commonly necessary smoothing of the spectra to reduce noise (e.g., from inefficient subtraction of atmospheric bands, acoustical or electrical fluctuations, interference fringing, etc.). Water is known to have vibrational transitions within this spectral region which, especially in the case of atmospheric water vapor, can introduce a large amount of “noise”. Water vapor absorption is reduced instrumentally by background subtraction scans; however, sharp features are still observed from varying humidity levels along the beam path. 32 Although purging with a dry gas improves the spectra, in most cases, complete elimination of water vapor absorption is either impossible or, if implemented, perturbs the native environment of the sample. Generally, interference from water vapor absorption is inevitable and hence the resulting spectra, more often than not, have significant contributions from variation in environmental humidity. Proteinaceous species are most commonly studied in the solid or liquid phase which leads to broad absorption bands. Atmospheric water vapor contributions appear as comparatively sharp peaks superimposed on the underlying substrate spectrum. Common approaches for dealing with this usually involve either smoothing of the data or application of an atmospheric correction.33,34 Both approaches have intrinsic characteristics that potentially can lead to spectral alteration. Dealing with these artifacts is also of great importance because deconvolution of the resulting spectra can be significantly affected as, consequently, will be any assignment or quantitative analysis of the contributing bands.
Traditionally, Fourier transform infrared (FT-IR) spectroscopy has been widely used to acquire spectra of proteinaceous substrates and hence has attracted much attention as a means of generating the best possible spectra yielding the most effective subsequent spectral analysis. However, with the development of other techniques such as photothermal atomic force microscope infrared spectroscopy (AFM-IR), sum frequency generation (SFG) spectroscopy, or infrared reflection absorption spectroscopy (IRRAS) and polarization IRRAS (PM-IRRAS), which can often yield very low spectral intensities, noise contributions are exacerbated and can arise from several sources, with different relative significance depending on the specific measurement technique. Many of these spectroscopic techniques are now commonly used to acquire protein spectra which can hence exhibit different noise features of varying magnitude and frequency. FT-IR benefits from the Fellgett advantage where intensities at different spectral wavenumbers are recorded simultaneously and hence there is no temporal variation within a given spectrum. This can result in the efficient subtraction of water vapor peaks from the measured spectrum when background subtraction is applied giving substantial improvements in the spectrum. However, the extent of the improvement is determined by the measured spectral resolution. The higher the resolution, the sharper the interfering water vapor bands. In turn this makes the determination of a subtraction coefficient and subsequent subtraction more accurate. 32 Nevertheless, there remains an uncertainty in the completeness of the subtraction and small residual water vapor contributions can remain. This can be of particular significance for substrates or techniques where the signal intensity is intrinsically low, with consequences for the quantitative analysis of the resulting spectra. Even for a spectrum recorded using FT-IR spectroscopy, where water vapor may be the only interfering contribution, smoothing is still widely used as a supplementary technique. This does, however, represent the ideal situation and it can often be the case that other sources of noise are present in the spectrum, e.g., periodic oscillations from interference fringing, and hence subtraction becomes insufficient at generating spectra suitable for quantitative analysis. 35 For other techniques where spectra are measured in finite times, such as in AFM-IR or SFG spectroscopy, there is an unavoidable temporal factor in the measurement.36,37 Other noise features such as high-frequency periodic oscillations, e.g., from acoustic or electrical interference, can then appear in the spectra and subtraction is no longer appropriate, and smoothing/filtering methods can be much more suitable and effective. This means that other forms of noise and atmospheric variations must be considered so that simple atmospheric subtraction alone is unlikely to be satisfactory. Thus, in all cases, including FT-IR, a more general approach for dealing with high-frequency noise is essential.
In this work, we will discuss the two common topics of deconvolution and spectral smoothing associated with the analysis of highly convoluted spectra which exhibit high-amplitude, high-frequency noise, a regular occurrence in the amide I band in proteinaceous substrates. The most frequently applied techniques and algorithms for analyzing and visually improving spectra are assessed for their benefits and associated limitations. In order to achieve this, both simplified models as well as actual spectra will be employed. From this analysis, our aim is to determine an improved method for spectral smoothing and subsequent quantitative analysis by means of peak identification and band deconvolution. This will lead to greater accuracy when analyzing proteinaceous substrates, a critically important outcome for biological applications. Finally, the quantitative comparison of the vibrational spectra of two or more distinct samples will be examined with particular emphasis on shifting baselines and how to account for them in order to arrive at a valid similarity index.
Signal Smoothing
Although spectral smoothing is widespread in many branches of spectroscopy, the algorithms in use are highly varied. 38 Their choice is often subjective and influenced by what produces the best visual result. Smoothing is typically used for two reasons: reducing or removing random instrumental noise and/or reducing or removing distinct sharp “peaks” in the spectra. The sharper the peaks, the more severe the smoothing as determined by the selected smoothing input parameters. All smoothing algorithms are intrinsically “lossy”, and excessive smoothing will result in a reduction of underlying spectral information and deviation from the “true” spectrum. Three of the most commonly used smoothing algorithms are discussed below with particular relevance to their application in the analysis of the amide I vibrational band and its concomitant water background.
Adjacent Averaging
Arguably, the simplest method is adjacent averaging (AA) of spectral data points. This method changes each data point to become the average from a given number of adjacent data points equally distributed either side (where spectral endpoints receive a truncated window) as shown in Eq. 1
Although simple, AA is often very effective giving visually smooth spectra without significant loss of the underlying spectral information. It does, however, only work best when variations in the intensity of adjacent points are small otherwise larger numbers of adjacent points are required and the loss of spectral information increases. A notable drawback, apparent in Figure 1a, is due to the algorithm tending towards a flat line with an increasing number of averaging points meaning spectral peaks and features always become less intense regardless of how sharp they may appear in the recorded spectrum. Sharper features are reduced more proportionately than broad ones hence this method is more appropriate for reducing noise of this type, typically expressing high frequency variation in the spectral data. Conversely, AA is, however, normally far from ideal for analysing the broader feature of the Amide-I band. Water vapour “spikes” are very sharp and often of very high intensity. Their removal usually requires averaging over a high number of points and therefore leads to significant loss of spectral information. This is likely to affect the underlying spectral data which has consequences for subsequent peak-fitting procedures and comprehensive peak analysis.
Identification of Peak Centers
An additional consideration for signal smoothing is the identification of peak centers in the resulting spectra. Typically, peak centers are identified using derivative spectra. Specifically, finding the local maxima in the spectrum, including partially resolved shouldered peaks, by progressive analysis of higher order derivative spectra, is discussed later. This is particularly relevant in the case of the amide I band where the contributing features are close together necessitating the use of higher order derivatives.24,25 Figure 1b shows the second derivatives of the band with different degrees of AA smoothing. As expected, for the least-smoothed spectra, the second derivatives show a spectral minimum at the peak center. For increased degrees of smoothing, however, the second-derivative spectra start to deviate from the pronounced peak center, eventually splitting into two. This is due to the smoothing window becoming increasingly wide compared to the peak width. This effect is also responsible for the flattening of the central part of the spectra shown in Fig. 1a. Clearly this indicates that excessive smoothing can lead to significant errors in determining the peak centers, number of contributing peaks, and their relative contributions.
Model Gaussian profile peak with a mean of 50 and standard deviation of 5. Adjacent averaging smoothing with (a) smoothed spectra with progressive increases in the number of adjacent points included within the averaging window (labeled as AAN where N is the number of adjacent points on either side, Eq. 1), (b) second derivatives of the smoothed spectra, (c) progressive smoothing of the same spectrum as in (a), but with superimposed high-frequency noise (modeled by discrete sinusoidal variations with an amplitude of 0.02 and frequency of 10), (d) second derivatives of the AA0, AA1, AA2, and AA3 smoothed spectra in (c), and (e) second derivatives of the AA4, AA5, AA10, and AA20 smoothing spectra in (c). The second-derivative plots of the high-frequency noise spectra have been separated into two plots for clarity. A nominal resolution of one unit per point is used for the spectral models.
An important feature of using derivative spectra for peak identification is that higher derivatives are more affected by sharp peaks (owing to their large gradients) and hence broader peaks are more susceptible to being obscured by those of water vapor or other high-frequency changes. This effect is demonstrated in Figs. 1c to 1e where high-frequency noise, modeled as discrete sinusoidal variations, has been superimposed on the underlying Gaussian profile. This noise could be due to water vapor or alternative high-frequency noise from acoustical, electrical, or other environmental effects (although it should be noted that water vapor changes would be expected to be present with a non-sinusoidal profile, discussed later). Although the peak may still be observable in zeroth order, its second-derivative spectra are obscured by the noise even at the highest degrees of smoothing. Therefore, for valid identification of spectral peaks in the amide I band, smoothing of the spectra must have eliminated as much of the high-frequency noise as possible otherwise the higher order derivatives will not reveal the true peak centers.
Savitzky–Golay Filter
Another common algorithm adopted for spectral smoothing is the Savitzky–Golay (SG) filter. This method utilizes sum of squares regression to fit polynomials of a given order through a given sub-set of adjacent data points in the spectrum.
39
The input parameters are: the polynomial order with which to fit the data and the number of points to include in the fit with a minimum of one greater than the polynomial order (so that a specific polynomial can be identified by the algorithm). The actual algorithm uses sets of weighted coefficients to determine the new, smoothed data points rather than actually undertaking polynomial fittings using sum of squares regression as shown in Eq. 2
Since this algorithm essentially fits a polynomial through the data points, it is significantly less perturbing than the AA method and hence results in smaller deviations from the underlying peak shape for the same degree of smoothing. It is, however, still a lossy method and therefore, with an increasing degree of smoothing, will lead to progressive reduction in spectral information. The effect of increasing the smoothing window width for this algorithm is illustrated in Figure 2 for both the underlying spectrum and its second derivative.
This SG method reassigns the value at each data point in the spectrum to be that of the fitting polynomial at that position. However, the generated polynomial is unique to each data point and hence the resulting smoothed spectra by no means have to be smooth in the mathematical sense. This, therefore, can result in sharp features in the smoothed spectra which can dominate the higher derivative spectra as discussed above. Additionally, any high-frequency contributions will often require significantly wider smoothing windows (strongly dependent on the amplitude of the noise owing to the method employing regression analysis). This can be observed in Fig. 2 where high-frequency noise has been superimposed in the same way as for the AA method analysis in Fig. 1. In the second-derivative spectra of the smoothed data, the high-frequency contributions are suppressed significantly with increasing smoothing window widths, showing much better results than the corresponding analyses using AA (Figs. 1d and 1e). This, therefore, shows how just moderate SG smoothing can visually reveal the underlying peaks in the second derivative. However, for higher derivatives, greater smoothing is required since the high-frequency noise has only been suppressed rather than completely eliminated because, as mentioned previously, the residual noise becomes more prominent with progressive differentiation. Hence, whilst the SG algorithm achieves effective smoothing, its application is primarily for visual improvement of the spectrum. When high-frequency random noise is present in the spectra, e.g., from water vapor, peak identification remains difficult even with high degrees of smoothing.
Model Gaussian profile peak with a mean of 50 and standard deviation of 5. Upper panels: Smoothed spectra using the quadratic Savitzky–Golay algorithm where the sequence shows progressive increases in the smoothing window width (SGN indicates N points are used to fit the quadratic function). Lower panels: Smoothing algorithms have been applied to the same spectra with added high-frequency noise (modeled with discrete sinusoidal variation with an amplitude of 0.02 and frequency of 10). A nominal resolution of one unit per point is used for the spectral models.
Fast Fourier Transform Filter
The fast Fourier transform (FFT) filter smoothing algorithm calculates the discrete Fourier transform (DFT) of the full spectrum across all frequencies rather than only considering a point window as in the two previously mentioned methods, as shown in Eq. 3
In the frequency domain, the spectral features and the noise ideally become separated. The next methodical step involves the application of a filter to the resulting DFT in order to affect the influence that each frequency has on the time domain spectrum. Typical filters include high- or low-pass filters, which commonly have either absolute (step function) or parabolic cut-offs, bandpass filters, or bandblock filters. 41 The choice of which filter to use is very case dependent and usually based on knowing the types of noise present in the data, i.e., high frequency, low frequency, etc. Spectra with high-frequency noise, which is common for vibrational IR spectra, will typically benefit most from a low pass filter. Once the filter is applied, thereby altering the frequency domain spectrum, an inverse Fourier transform (IFT) is applied to convert back to the time domain followed by re-insertion of the baseline to finally yield the smoothed spectrum.
Although this algorithm is also technically lossy, the loss in spectral information is mostly specific to certain frequencies, depending on the choice of filter. Hence, in the case of an IR spectrum with high-frequency noise, and where the application of a low pass filter is appropriate, the noise and genuine spectral features become segregated in the frequency domain. If an appropriate filter cut-off frequency is chosen, then the FFT filter smoothing will lead to complete elimination of the spectral noise and minimal deterioration of the underlying spectrum. A demonstration of the application of the FFT filter method is shown in Figure 3 where sinusoidal noise at two frequencies has been applied to modelled Gaussian peaks of different widths in the time domain and their FFT calculated (shown in the third row). The indicated threshold frequency is applied just below that of the lower of the two noise frequencies. It is clear that narrower true peaks in the time domain (i.e., those with peak widths more similar to that of the noise oscillations) have a broader distribution in the frequency domain. Hence the FFT filter removes some spectral information and subsequently there is a greater deviation from the true spectrum when converted back to the time domain. Conversely, for sufficiently separated spectral peaks and noise in the frequency domain, the filter results in no observable deviation in the spectral peak whilst there is complete elimination of the noise.
Gaussian and Lorentzian Line Profiles
In addition to considering the application and effect of filtering on the line frequencies, it is also necessary to consider its effect on the line profiles. For most cases, a Voigt profile provides the best model since it combines both Gaussian and Lorentzian functions by convolution and hence by addition in the frequency domain after application of a Fourier transform.
42
In either case, however, the Fourier transform of Gaussian or Lorentzian functions yields infinite width functions as shown in Eq. 4
Because of their infinite width, the assertion that the high-frequency noise can be separated from the spectral peaks cannot be completely accurate. Therefore, a filter to remove high-frequency noise components always results in some loss of spectral information. This effect is however usually small given that the high-frequency noise is significantly separated from most of the spectral components in the frequency domain. Therefore, for intermediate smoothing by setting moderate-to-high threshold frequencies on the applied low pass filter, the modification of the underlying spectral peaks is minimal. Nevertheless, when increased smoothing is applied by decreasing the filter threshold, larger deviations can occur. The effect of this can be seen in Figure 4 where a modelled Gaussian is subjected to decreasing threshold low pass filters.
Figure 4 also shows that, unlike the smoothing methods previously discussed, the FFT filter succeeds in completely removing the high-frequency noise component once the threshold frequency surpasses that of the noise. This is especially evident in the second-derivative spectra which demonstrates the importance of second-derivative spectra in peak-finding calculations. The difference spectra in Figure 4 show the different degrees of smoothing compared to the true underlying peak intensity as well as the equivalent for the second-derivative. Once the appropriate threshold has been reached, the deviation from the true peak are minimal however it is also clear that further smoothing increases the loss of spectral information. Nevertheless, this algorithm is significantly better for minimising the loss of spectral information whilst removing spectral noise and leading to high-confidence in determining peak centers.
It should also be noted that water vapor bands in the spectrum should appear with Lorentzian or Gaussian profile, i.e., would have infinite width in the Fourier domain. Therefore, the previous analysis of high-frequency noise at first sight appears not to be directly applicable to water vapor lines but only to sinusoidal oscillations in the spectra, e.g., interference effects in FT-IR spectra, or acoustical or electrical noise in spectra recorded in a step-wise fashion in a finite time for example using AFM-IR spectra. Nevertheless, as with any mathematical function, the variations from water vapor bands can be assigned an infinite Fourier series and hence approximated by a finite set of sinusoids. The above treatment can therefore remove significant high-frequency contributions from the Fourier series and minimize their contributions to the spectra. This effect is also apparent by considering the Fourier transform of the sharp water vapor bands that are converted to broad contributions in the Fourier domain as demonstrated by Eqs. 4 and 5. Therefore, an appropriately chosen filter can remove a significant proportion of the water vapor contribution in the frequency domain, leaving the genuine spectral features, present as significantly narrower bands, virtually unperturbed when converted back to the time domain. To conclude, although the above treatment of sinusoids may appear theoretical in relation to real spectral noise, these periodic variations can often be found in spectra and a similar treatment can be used to substantially reduce, though not completely eliminate, the water vapor contribution.
Doublet Peak with Noise
In order to implement spectral deconvolution, overlapping bands must be able to be distinguished and hence the chosen smoothing algorithm is required to sufficiently remove any high-frequency noise in the spectral data whilst maintaining the contributing peak information. In order to further demonstrate the effectiveness of the FFT filter method compared to the AA or SG algorithms, all three were applied to a model spectrum consisting of two close, partially resolved, Gaussian peaks with large amplitude noise superimposed on them (modeled by discrete sinusoidal variation) as shown in Fig. 5. Additionally, increasing degrees of smoothing using each method previously described are shown alongside their second derivatives and difference spectra compared to the true underlying spectrum. The difference spectra show that all methods lead to loss of information which increases with the degree of smoothing; however, the FFT filter is the only method which completely removes the high-frequency noise contribution and still manages to achieve minimal loss of signal.
Modeled Gaussian peak with applied sinusoidal noise of amplitudes 0.7 and 1.0 at frequencies of 1.1 and 1.6 (as shown in the frequency domain). Shown in rows from top to bottom are: The true Gaussian peak, combined Gaussian and noise, FFT of the noisy spectrum with imposed frequency threshold (red dashed line), IFT of the post filter frequency domain spectrum, and the calculated difference spectrum to the true Gaussian as in the top row. Columns (a) to (c) show identical treatments with decreasing width Gaussian peaks in the time domain, i.e., increasing similarity of the true peak width and that of the noise. The Gaussians modeled are (a) amplitude of 50 and standard deviation 1, (b) amplitude of 20 and standard deviation 0.4, and (c) amplitude of 10 and standard deviation 0.2, all with a mean of 5. A nominal resolution of 0.02 units per point is used for the spectral models. Modeled Gaussian peak with a mean of 50 and standard deviation of 5. Left-hand panels: Zeroth-order spectra, right-hand panels: Second-derivative spectra. Top: Noiseless smoothed spectra using an FFT filter. Middle and bottom: Progressive decreases in the low-pass filter threshold frequency (FFTN indicates a threshold frequency of 1/N). Additionally, the smoothing filters have been applied to the same spectra with added high-frequency noise (modeled with discretized sinusoidal variation with an amplitude of 0.02 and frequency of 10). Bottom: Shown are the difference spectra compared to the underlying true spectrum. A nominal resolution of one unit per point is used for the spectral models. Modeled IR spectrum showing two Gaussian peaks centered at positions 40 and 60 with standard deviations of 5. Superimposed on the spectrum is high-amplitude, high-frequency noise (modeled with discrete sinusoidal variation with an amplitude of 0.02 and frequency of 10). The smoothed spectra, subsequently calculated second derivative spectra and difference spectra, calculated from subtraction of the true spectrum from the smoothed spectra, are shown for the AA, SG, and FFT filter algorithms with varying degrees of smoothing from 1 to 10 points for their respective input parameters. A nominal resolution of one unit per point is used for the spectral models.


Practically, the noise will likely be a combination of variations presenting at a range of oscillation frequencies in the spectrum, most of which will present as high-frequency spikes relative to the underlying bands. Hence, the FFT filter, with an appropriately chosen threshold frequency, should remove all high-frequency components leaving minimal noise in the spectrum without significantly degrading the spectral information. Based on the analysis given above, it is clear that to deconvolute the amide I band for protein structure determination, FFT smoothing is the method of choice.
Derivative Analysis for Peak Fitting
As mentioned above, spectral analysis commonly requires the identification of peak centers and the subsequent integration of the complete peak profile so that their relative contributions to the total spectral intensity can be quantified and assigned. Typically, peaks that are well resolved so as not to overlap significantly and distort the profiles present with clear local maxima within the band where there is minimal deviation in the wavenumber of the local maxima within the band and that of the underlying peak contributions. Therefore, contributing peaks can be identified either by finding local maxima in the spectra or zeros in the first derivative spectrum where the second derivative is negative. As mentioned earlier, the derivative spectra, although being trivial to calculate, are increasingly susceptible to high-frequency noise. Because this effect increases the higher the derivative order, identification of peaks by probing the lowest order derivatives is usually best. For the amide I band, where water vapor is prevalent and often hard to remove entirely from the spectra, it is better not to use higher derivatives.
Where peaks are only partly resolved and have significant overlap that distorts the spectrum, peak centers cannot be easily identified from visual inspection of maxima in the spectra. In this case, higher-order derivatives are required to resolve the peaks since increasingly higher-order derivatives produce a reduction in the peak widths compared to the width in the zero-order spectrum. This is demonstrated in Figure 6a, where a single Gaussian profile has been used to model an IR peak and the corresponding zero to fourth-order derivatives are shown. Figure 6a also shows the characteristic shapes of the derivative spectra associated with a Gaussian profile peak where important features can be identified, namely: peak centres appear with zero amplitude in odd-order derivatives, and are at maxima for derivatives which are 0 mod 4, and minima for derivatives that are 2 mod 4. Based on these observations, it might be concluded that any order derivative could be used to identify peak centers. However, on the introduction of a second peak in intermediate proximity to the first, as shown in Figure 6b, the peak centers are not immediately identifiable from the zero- or first-order derivative spectra. Only on further differentiation to reach the second and higher derivatives are the peaks distinguishable. A complication that arises when using derivatives of order three or greater is that the defining feature that is aligned with the peak centre, i.e., zero, maximum or minimum, is not unique to the peak centre. This is clear in Figure 6a where the third-order derivative shows three zeros associated with the underlying peak and the fourth-order derivative shows three maxima. This ambiguity can however be circumvented by using a combination of derivatives, i.e., first and third, or second and fourth. When pairs of derivatives are used, then the analysis also benefits from the greater resolution of the peak centers from the higher-order derivatives without misidentifying peak centers from the additional, satellite spectral features. Another example of overlapping peaks is shown in Figure 6c, where now the peaks are sufficiently close together as to only become resolvable in the fourth-derivative spectrum. This progression clearly shows how increasingly higher-order derivatives are required for peak centers that become closer together (relative to their standard deviation).
For the amide I band, due to the closeness of the underlying spectral peaks, the result is a broad band with peak centers that often cannot be distinguished from visual inspection. In this case, therefore, use of higher order derivatives is essential. This is in contradiction to the conclusion that lower order derivatives are more suitable for removing water vapor interference owing to the sharper nature of these peaks dominating higher derivatives. This emphasizes again the importance of careful selection of the smoothing algorithm: It needs to remove as much high-frequency noise as possible without significantly altering the underlying spectral structure.
Application to Real Spectra
In order to assess the inherent capability of each of the methodologies for smoothing and subsequent deconvolution into underlying peak contributions in the case of a real spectrum, each algorithm has been applied to an actual amide I infrared band of a biological substrate known to contain a large variety of different contributing proteins, namely hair, where spectra were recorded using AFM-IR spectroscopy on a cross-section of European, brown, human hair fiber. The resulting smoothed spectra and their second derivatives are shown (without co-averaging) in Fig. 7 for AA (column 1), SG (column 2), and FFT (column 3). No dry purge gas was used so that narrow, intense water vapor bands are significant features of the spectrum providing a test of the methodologies for implementing effective noise reduction. All three methods yield efficient atmospheric water vapor subtraction (second row panels with highest three smoothing regimes) with the FFT filter providing noticeably better smoothing than the others. However, it is clear that whilst the AA and SG methods yield cosmetically improved spectra, their second derivatives are ineffective for identifying the underlying peak contributions. The positions and frequencies of these contributions are indicated in the third and fourth row of panels as identified using the seven-point FFT filter.
The peak centers identified by this method and their assignment are given in Table I, where it is clear that they align with the established values given in the literature. It was found that, even for very large smoothing windows, both the AA and SG techniques still retain sharp features in their second derivative spectra (bottom panels). Consequently, there is significant loss in the underlying spectral information and a high likelihood of errors in peak identification even if they are resolvable in the second derivative spectra. On the other hand, the FFT filter method shows that even for intermediate smoothing, once the frequency filter threshold surpasses that of the noise, smooth second derivative spectra are observed with the elimination of the water vapor bands (bottom FFT panel). Once the noise is eliminated, the second derivative spectrum shows a comparatively flat trace, indicated by the much-reduced scale on the ordinate axis, where only the real amide bands contribute. A necessary proviso, however, is that it is apparent that excessive FFT smoothing can lead to significant loss in spectral profile information. This manifests itself as shifts in the underlying peak centers, peak intensity deviations and merging of adjacent peaks into broader, unresolvable bands. From this application of each of the methodologies discussed earlier to a real spectrum, it is clear that the models used previously provide a sound basis for analyzing the effects of different methods on actual convoluted spectra. To conclude: overall the FFT filter is the best analytical method but with the caveat that careful consideration has to be given to deciding the cut-off frequency.
Derivative analysis of modeled IR peaks showing top to bottom panels: The zero- (red), first- (orange), second- (yellow), third- (green), and fourth- (blue) order derivatives, respectively. (a) A single peak using a Gaussian of mean 50 and standard deviation 10, (b) two peaks moderately close together with means of 40 and 60 and both with a standard deviation of 10, and (c) two peaks but now in closer proximity with means of 43 and 57 and both with a standard deviation of 10. A nominal resolution of one unit per point is used for the spectral models. Actual amide I band of a proteinaceous biological substrate, European, brown, human hair, recorded without co-averaging and without a dry purging atmosphere using AFM-IR spectroscopy. The raw spectrum shows significant overlap of contributing bands and high-amplitude water vapor noise (first row of panels). Smoothed spectra and their subsequently calculated second derivatives are shown for the adjacent averaging, Savitzky–Golay, and FFT filter algorithms with varying degrees of smoothing as indicated in the legends: AAN represents adjacent averaging with 2N + 1 points in the smoothing window, SGN represents Savitzky–Golay smoothing with N points fitted with a quadratic function, FFTN represents an FFT filter with a threshold frequency of 1/N. Top row of panels: Raw spectrum and spectra with the two lowest smoothing regimes applied to the raw data, second row: spectra with highest three smoothing regimes, third row: second derivatives of top row of spectra, bottom row: second derivatives of spectra shown in the second row of panels. Also indicated on the derivative spectra in the bottom two rows of spectra are the peak positions determined from the seven-point FFT filter smoothing algorithm (the method with the least smoothing) to show elimination of water vapor spikes. The spectrum was recorded at a resolution of 1 cm–1 per point.

Similarity Index
Another type of spectral analysis that is commonly undertaken is the use of a similarity index as a method of comparing two or more spectra to elucidate their underlying similarity.43,44 Often, this is a qualitative exercise where a visual comparison is used to make assignments or justifications for the similarity of two or more species.
45
This may be sufficient for many applications but in other cases it may require a more quantitative approach. In these cases, some of the more popular methods include calculation of difference spectra from normalized spectra, which is also commonly associated with making qualitative comparisons. Subsequently, similarity indices such as sum of squares calculations as in Eq. 6
A recurring issue with these methods, however, is the assumption of a constant baseline for each of the spectra being compared. It is often the case, however, that spectrometers undergo an environmentally induced baseline shift due to a change in temperature, humidity, air pressure, etc. As is obvious from Eqs. 6 and 7, a shift in the baseline is automatically intrinsic to the calculation and will therefore affect the similarity result. This imposes significantly larger uncertainties on the quantitative comparison of two spectra and is difficult to quantify without knowledge of how the baseline shifts with the given environmental variation for a specific spectrometer. Therefore, a more quantitative similarity calculation is one which can eliminate as much of any baseline shift as possible from the calculations.
One obvious method is to adapt the sum of squares difference method whereby the sum of squares is calculated relative to a trend line rather than to the x-axis, thereby treating the trend line as the baseline shift and removing it from the calculation. A slight drawback arises from using normalized spectra, however, since the calculated normalization will include the baseline contribution. Therefore, if the baseline is to be excluded from the similarity calculation without the requirement to manually calculate and subtract it from each spectrum prior to the analysis, the resulting calculation must account for the normalization without the baseline shift contribution. One method of achieving this is to subtract the baseline difference from the first spectrum (i.e., the one being subtracted from rather than that being subtracted), thereby effectively putting both spectra on the baseline for the second spectrum (i.e., the one being subtracted). This does not remove the baseline contribution from the normalization, but it does apply an equal contribution from the baseline to each spectrum and hence the normalization should be equalized. This method is expressed mathematically in Eq. 8
A further allowance must be made where the spectra being compared have different intensities, which are not inherent to the spectra but associated with the measurement procedures. Clearly, this is a problem if S1 is significantly more intense than S2 despite being exactly the same spectra (e.g., consider Sim(10S1, S2)). This issue results in the difference spectrum not accounting for this shift in intensity and is rectified by shifting the normalization procedure to after this calculation to account for the baseline shift. A method of implementing this is to calculate the above similarity for varied multiples of one spectrum and take the minimum value as expressed in Eq. 9
An alternative approach comes from adapting the cross-correlation method. Considering each spectrum to be the sum of a true spectrum,
The final item that needs to be addressed is normalization since the result of the similarity calculation is heavily dependent on the intensities of the two spectra. A slightly different normalization method must be applied to derivative spectra since, unlike the zero order spectra, they are not always bound to be positive. Applying a simple shift so that the spectra fall between 0 and 1 is not valid. Instead, the relative heights of the derivative peaks should be consistent, in an analogous way to the zero-level being consistent in the spectra, once the polynomial baseline has vanished. This, therefore, allows the derivative spectra to be normalized by fixing the zero-level and expanding the y-axis such that the largest peak falls to either 1 or –1 (depending on whether it is a maximum or minimum). If this approach is applied to the derivatives of the same spectra, only with differing intensities, then it should yield the same result after applying the constant factor rule of differentiation as expressed in Eq. 13
As an illustration of how both techniques can be used to find the similarity index, as well as the configurations adopted to allow for a baseline shift, is presented in Fig. 8. Here two spectra, modeled by Gaussian functions where one of them features an intensity shift as well as a quadratic baseline (Fig. 8a), are compared and the similarity indices generated using each of the methods described (shown in Fig. 8c on a logarithmic scale for clarity). Modeled spectra were used to enable the calculated similarity indices to be accurately compared with expected values. (Such a benchmark comparison would not be possible for real spectra due to their inherent variability and subsequent uncertainty.) The two commonly used methods of sum of squares normalized difference spectra, and the cross-correlation method clearly give worse errors in the similarity index, expected to be 1, highlighting the significance that the baseline shift has on the calculations. The difference spectrum method, which makes the same calculation but allows for a non-zero polynomial trend line, clearly yields the expected improvement. This method will always yield a result at least as good as the original method but there is still an underlying error due to the nonlinear behavior of the baseline. On the other hand, when used with a quadratic trend line, as shown, a perfect result is obtained, suggesting further improvement can be achieved by fitting increasing order polynomial baselines.
Illustration of similarity index calculations applied to two spectra, with two absorptions modeled by Gaussians, one featuring a quadratic baseline shift. (a) The two spectra being compared, one with two peaks defined by 0.4 and 0.6 amplitude Gaussian functions centered at 30 and 70, respectively, both featuring standard deviations of 10 (black trace); and a second spectrum (red trace) which is identical but featuring a 1.3 factor intensity variation and quadratic baseline defined by the quadratic constants: a = 2E-6, b = –3E-5, and c = 0.007. (b) The third-order derivatives (unnormalized) of both spectra as used by the adapted cross-correlation method. The bar chart in (c) shows the error in each method used to calculate the similarity index (expected to be 1), plotted on a logarithmic scale for clarity (the final method column gives 0 error hence is absent from the plot). A nominal resolution of one unit per point is used for the spectral models. Amide I contributions identified in the spectrum presented in Fig. 7 using a seven-point FFT filter smoothing algorithm and subsequent second-derivative analysis.
Although real baseline shifts will not necessarily conform to linear or quadratic functions, increasing power regression methods will always yield at least as low an apparent error in the similarity index. It is worth noting, however, that should the order of the trend line be increased too significantly, contributions from spectral features rather than the baseline could be fitted leading to erroneous results. Nevertheless, when the calculation is restricted to reasonably low-order trend lines, i.e., approximating the baseline shift by a low-order polynomial, an improved and more accurate comparison of spectra will result. In this case, the cross-correlation method yields a better result than the difference spectra method, however, but still shows a significant inaccuracy. For the method adopted here, where third-order derivatives are used (shown in Fig. 8b), a much better result is obtained. Given the quadratic baseline, it would be expected that the similarity index should yield a perfect result. However, calculation errors are introduced due to using derivatives of discrete datasets and an underlying error is always likely to occur. In this case, the error is 3 parts per million (ppm) and minute, and hence would be expected to be similarly low in reality. This again emphasizes the ideal, theoretical case of a trend line that can be perfectly fitted with a quadratic function, although in real spectra, this idealistic case is unlikely and hence a perfect result should not be expected. Given the baseline shifts in actual spectra are generally broad, smooth functions, increasing order derivatives of the spectra will minimize the baseline contribution and hence yield improvements in the calculation with the caveat that utilizing derivatives of increasing orders will introduce increasing errors due to the discrete nature of the spectra. As demonstrated here if the baseline can be approximated by a relatively low-order polynomial shift, the errors due to the derivative calculation should be negligible. One final comment about this method follows from the earlier discussions about smoothing and derivative analysis where the high-frequency noise becomes more apparent in higher derivatives. As this calculation uses relatively high-order derivatives, the requirement for significant noise reduction is clear.
Discussion
The essential elements involved in analyzing spectra such as the amide I band are the ability to smooth the data in order to remove or reduce the influence of random noise, being able to deconvolute the spectrum into its contributing peaks and allowing two spectra to be quantitatively compared for similarity. The advantages and disadvantages of three of the most common smoothing algorithms used for IR spectra are discussed here, with particular reference to their application to the amide I band. The simplest and easiest to apply is the AA method. This unsophisticated smoothing algorithm works best for spectra with smaller amplitude variations. The amide I band, which contains very sharp water vapor peaks, often requires a very wide spectral window for sufficient smoothing to visually remove them. The consequence of applying a widening averaging window is that in the limit the spectrum tends towards a flat line and, as demonstrated in the model spectra in Figs. 1c to 1e, even with high smoothing the high-frequency noise still remains to some extent. Additionally, significant deviations from the underlying spectrum occur when the smoothing window becomes too wide (i.e., significantly wider than the standard deviation of the peak) potentially leading to errors in the subsequent peak analysis. A second consequence is that high levels of smoothing begin to distort the underlying spectral shape, thereby making it impossible to confidently deconvolute the spectrum. Although this may not be important in the zeroth order spectrum, it will dominate the higher derivative spectra commonly used for peak identification. When applied to the amide I band, the AA smoothing algorithm should be considered as a cosmetic method providing little assistance in reducing noise prior to subsequent deconvolution.
The SG method essentially performs sum of squares regression to fit an nth-order polynomial through a spectral window of defined width. It is considerably less severe than AA and can also lead to significant visual improvement to the spectra. Once again, however, due to the unique nature of the set of data points used in the smoothing window leading to unique polynomials for each data point, the resulting spectrum often still retains sharp features which will adversely affect the higher derivative spectra. Hence, although this method may be effective in visually producing a good reduction in the noise, in the case of the amide I band it causes difficulties in the use of higher derivatives for deconvolution. Therefore, this method, such as AA, is primarily cosmetic.
The FFT filter method at first calculates a DFT of the data points which converts the spectrum into its contributing frequencies in the frequency domain. Depending on the extent of overlap of the spectral and noise contributions in the frequency domain, by applying an appropriate filter to the noisy spectrum, then in principle elimination of the noise can be achieved with minimal loss in the underlying spectral shape. In the case of the amide I band, it results in visually smooth higher derivative spectra with the absence of the water vapor peaks that were previously dominant features. Hence, the FFT filter is an excellent method for application to the amide I band where higher order derivatives are required for confident deconvolution and to reveal the fundamental peak contributions. The only limitation that arises is whether there is incomplete segregation of the frequencies of the spectrum and the noise in the frequency domain leading to some loss in spectral information. This has the effect of reducing the resolution of underlying peaks contributing to the band that are within this spectral width. Thus, should the chosen cut-off frequency result in complete elimination of all frequencies above 1/n, corresponding to an n-point filter, then any underlying bands that also lie within this limit of n points will not be resolved. They will need to be fitted and treated as a single, combined peak. Although this is a limitation of the FFT method, it nevertheless provides reliable assignments of the spectral contributions and is clearly the best choice for analyzing amide I spectra.
The second topic relevant to analyzing the amide I band is deconvolution and peak identification leading to the identification of the underlying contributions from different protein secondary structures. Generally, spectra with minimal overlap of peaks are easily assigned with their recorded positions virtually unchanged from the underlying peak centers. Where this is not the case, however, identification of peak centers requires higher derivatives to be examined. The closer together the peaks are, the higher the order of derivative needed to resolve them. Overlapping is sufficiently common in the amide I band that a generally useful method is to analyze the second derivative spectra to find all the minima which correspond to the peak origins. However, in some cases this is not sufficient and higher order derivatives are required to resolve peaks. A difficulty that arises is that the higher derivatives have satellite spectral features which can be mistaken for peak contributions. This dilemma is resolved by combining the use of the second and fourth derivatives, whereby the fourth derivative is used to find all minima in the second derivatives. This not only circumvents the ambiguity of additional features in the higher derivatives but also achieves better resolution of peak centers.
Once the peak centers are unambiguously identified by this method, then the spectrum can be fitted using these contributions by simulated annealing of their peak heights, peak centers, and standard deviations, i.e., full width at half-height (FWHH) values. Generally, this method of simulated annealing determines the local turning point in the output parameter which, in this case, is the minimum of the sum of squares regression for the difference spectrum between the fitted spectrum and actual spectrum. It is common practice, however, to apply boundary conditions on the contributing peaks such as restricting their peak centers and standard deviations and/or discarding negative peaks. These boundary conditions are chosen so that the identified peak centers align with the true, deconvoluted spectrum and all peaks having non-negative intensities and standard deviations within defined limits, i.e., based on how sharp the band is expected to be. The peak centers themselves are identified using the derivative analysis described above and may be slightly shifted from their true values due to overlap with other peaks. This shift in the peak center will be reduced in the higher derivative spectra and therefore, by using a combination of second- and fourth-derivative spectra, greater confidence can be placed in the peak centers. Boundary conditions placed on the standard deviation of the peak (FWHH) should be less severe than on the other variable parameters such as the peak centers and be used to prevent excessively sharp or excessively broad peaks from being fitted in the spectrum. However, the boundary conditions used in practice will be dependent on the intensity of the specific peak as well as the type of environment that the peak is assigned to, i.e., it is expected to arise from a species where a large variation in environment might occur leading to a broad band such as one involving H-bonding. A summary of the suggested procedure to achieve high-confidence deconvolution is demonstrated in the flow chart in the Supplemental Material (Figure S1).
The final topic considered here for a thorough spectral analysis has been the quantitative comparison of two or more spectra in order to generate a similarity index. Although much work has already been reported on this topic, the methods used do not generally allow for a shift in the baseline between two spectra, a common feature of most spectrometers because they are highly environmentally sensitive. Here, we have considered the modification of two common methods used for quantitative comparison, which do not intrinsically allow for a baseline shift. The two methods examined, the sum of squares regression of difference spectra and the cross-correlation methods, are modified using similar algorithms to account for a shift in baseline. The first modification, based on the sum of squares regression, uses the sum of squares from a calculated trend line of the difference spectrum rather than from the x-axis and hence this allows for a shift in the baseline which can then be fitted with a selected type of trend line, i.e., linear, spline, etc. Complications can arise with this modification when considering the effect of unnormalized spectra. Normalization is executed post-calculation of the difference spectrum and implemented by dividing by the sum of each spectrum having put both spectra on the same baseline by subtracting the calculated trend line from one of the spectra. Moving the normalization to the end, however, no longer accounts for variation in overall spectral intensities. To circumvent this difficulty, the similarity calculation is repeated for different multiples of one of the two spectra and then extracting the minimum value, which will correspond to the closest subtraction. The second common method to generate a similarity index applies the cross-correlation algorithm to both normalized spectra, which produces greater outputs for spectra that are more similar. When a difference in the baseline is introduced, however, then the cross-correlation procedure becomes complicated with many cross terms involving undefined baselines. To circumvent this problem, the baselines can be fitted with an nth-order polynomial where subsequently calculating the cross-correlation of the (n + 1)-order derivative therefore eliminates the baseline from the spectrum and yields the desired comparison.
Conclusion
Here, we have critically examined the available methods of smoothing, deconvolution, and fitting, as well as the quantitative comparison of spectra, for the quantitative analysis of the amide I band of proteinaceous substrates. The benefits and limitations of using the most commonly employed algorithms are considered, and these have been modified as necessary to enable them to be applied accurately. Although the AA and SG smoothing algorithms are in common use, it is concluded that when used to smooth heavily convoluted spectra with high-frequency noise having significant amplitude, as is often the case for the amide I band, the result is mainly cosmetic. Nevertheless, it is shown that they are fully capable of producing significantly improved second derivative spectra, but these remain dominated by the high-frequency noise even at large widths of the smoothing window. Sufficient elimination of the noise in the second-derivative spectra is only achieved, if at all possible, by imposing excessively high smoothing with subsequent significant loss in spectral information. On the other hand, it is found that smoothing with an FFT filter not only improves the appearance of the spectra but also eliminates high-frequency noise components without significant loss in spectral information. Using both model and real spectra, it is demonstrated that once the low-pass threshold frequency surpasses the noise frequencies, the effect of noise on the spectra and on their derivatives is removed leaving second-derivative spectra where the underlying spectral information is retained. An important caveat, however, is that the application of a frequency filter always results in some loss in spectral information. The Fourier transform of Gaussian and Lorentzian functions, both having infinite wings, requires critical selection of the filter threshold, since excessive smoothing will lead to significant deviations from the real spectrum and consequently reduced confidence in identifying peak centers. It is shown that a second limitation of the FFT smoothing filter is that any filter cut-off frequency results in loss in spectral resolution below the threshold. This can be rectified in the subsequent analysis by assuming that any bands that lie closer together than the filter threshold all contribute to the identified peak. However, their relative amplitude contributions cannot be quantified. Nevertheless, the FFT filter provides the best spectral smoothing for highly convoluted spectra that also have high-amplitude, high-frequency noise. After effective FFT smoothing, the combination of second- and fourth-derivative spectra yield robust identification of the contributing peak centers and an analytically valid spectral assignment. A further relevant aspect of spectral analysis considered here is the generation of a similarity index. A shortcoming in many methods is that shifts in the baseline are included in the spectra but not allowed for the comparison, thereby giving rise to erroneous results. In the present work, we have applied modifications to two frequently used similarity indices, namely the sum of squares regression of difference spectra and the cross-correlation analysis methods, to correct for any differences in the baselines of the input spectra.
Supplemental Material
ASP898536 Supplemental Material - Supplemental material for Spectral Analysis and Deconvolution of the Amide I Band of Proteins Presenting with High-Frequency Noise and Baseline Shifts
Supplemental material, ASP898536 Supplemental Material for Spectral Analysis and Deconvolution of the Amide I Band of Proteins Presenting with High-Frequency Noise and Baseline Shifts by Alexander P. Fellows, Mike T.L. Casford and Paul B. Davies in Applied Spectroscopy
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
We are grateful to Unilever R&D and the EPSRC for providing an iCASE studentship for A. P. F. on Grant EP/R511870/1.
Supplemental material
The supplemental material mentioned in the text, consisting of Figure S1 showing the flow chart summarizing a suggested procedure to achieve high-confidence spectral deconvolution, is available in the online version of the journal.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
