Abstract
Lupins (Lupinus angustifolius) are a type of pulse known for their high protein content, nutritional benefits and versatile uses in food. Lupins have a characteristic thick outer hull which may impact the development and application of NIR (near infrared) calibrations for assessing compositional traits. This study seeks to compare the predictive abilities of models based on wholegrain and ground grist for key quality traits, specifically protein and moisture content. To determine if the thick outer hull poses an obstacle for NIR light to penetrate and reflect from the seed, NIR spectroscopy of lupin samples was conducted for both wholegrain and ground grist forms. The same set of 120 samples was used to construct two prediction models (with a calibration set of 96 samples and an independent validation set of 24 samples). The results show that the predictions for moisture content were comparable between wholegrain and grist for the validation set, achieving r2 values of 0.99 for both models. The models for protein content exhibited a greater distinction between wholegrain and grist, yielding r2 values of 0.91 and 0.96, respectively. Overall, both sets of models exhibited high correlations with protein and moisture constituents and are indicative that calibrations developed on intact lupin hull using wholegrain NIR spectra can be used to predict protein content. Despite a decline in accuracy when compared with the grist results, the benefit of applying the wholegrain calibration outweighs the time taken to prepare the grist samples for NIR analysis.
Introduction
Lupinus angustifolius, commonly known as the Australian sweet lupin or narrowleaf lupin, has become increasingly recognised in recent years due to its notable nutritional profile and versatility in food applications. Lupins and other pulses high in protein content are becoming increasingly sought after as viable protein-rich alternatives to meet dietary needs. 1
Australian sweet lupin seeds have high protein content, dietary fibre, essential unsaturated fatty acids, and are low in starch content. 2 The sweet lupin is generally low in alkaloid content resulting in reduced bitterness compared with other lupin species. 3 Australian lupin production reached ∼1.3 million tonnes in 2023; an increase of ∼33% from 2022. 4 The increase in lupin cultivation has emphasised the importance of establishing efficient and reliable methods to assess lupin composition, particularly in large-scale breeding programs and production areas. Near-infrared (NIR) spectroscopy is ideal for routine analysis of grain quality traits, with several advantages such as being non-destructive, rapid, and capable of handling high throughput. This capability supports breeders in identifying optimal breeding lines that exhibit higher protein content. 5 Previous studies have explored NIR scanning of whole seed lupins and lupin kernel meal for various compositional traits, yielding varied results.6–9
Lupins are characterized by a thick hull or seed coat, and this could pose a challenge for NIR analysis. The lupin seed coat represents on average 24% of the seed weight, and seed coat thickness can vary between ∼257 and 335 µm, dependent on variety and environmental conditions.10,11 The thick hull, and high proportion of hull to cotyledon, may affect the penetration depth of NIR light which is dependent on multiple factors such as the physical proportions of the seed. 12 This could potentially lead to variations in the accuracy of lupin wholegrain predictions, especially as the seed coat only contains 3–4% protein content. 13 This study aims to explore whether the intact lupin hull has a negative impact on NIR predictive accuracy for quality traits, protein and moisture content, by constructing two calibration models, intact wholegrain and ground grist, and comparing the statistical performance of predicting an independent validation set.
Methods
Sample processing
A range of Australian lupin germplasm were acquired from lupin historical trials, commercially released varieties, spanning across harvest years 2015, 2017, 2020, and 2023 were used for calibration construction. The samples were sourced from the primary sweet lupin production regions in Victoria and South Australia. These samples, which were harvested in the dry season, were stored in screw lid plastic containers at ambient room temperature (∼16–21°C) in laboratory seed store. NIR spectra were acquired using a XDS Rapid Content Analyzer (FOSS, Hillerød, Denmark), which employs diffuse reflectance to capture wavelengths from 400 to 2498 nm. Prior to NIR spectroscopy, samples were moved to the same room as the XDS and let to equilibrate at room temperature (21 ± 1°C) for 24 h before analysis. Wholegrain scans were collected in triplicate, from ∼50 to 60 g samples, using a rectangular cell (repacked between scans). 14 Subsequently, lupins were subsampled and ground into grist using a Laboratory Mill 3303 (Perten Instruments, Stockholm, Sweden). Particle size distribution was determined via mechanical sieving (AACC International Method 55–06.01), and found to average 238 µm, with the largest particles of 280 µm. 15 The grist was then packed into ring cup cells (∼5 g sample) and scanned in triplicate (repacked between scans). Following this, lupin grist was analysed for moisture and protein (dry basis) content on the same day, with each sample run in duplicate. Moisture content was determined according to Moisture-Air Oven Methods (AACC International Method 44–15.02) via thermogravimetric analysis (TGM 8000, LECO Corporation, Saint Joseph, USA). 15 For protein content determination, the Crude Protein-Combustion Method (AACC International Method 46–30.01) was modified to use a furnace operating at 1200°C with a nitrogen analyser (FP928, LECO Corporation, Saint Joseph, USA). 15
Calibration development
WinISI software (v 4.6.11.14874, FOSS, Hillerød, Denmark) was used for calibration construction. The calibration set (n = 96) for model development consisted of samples grown prior to 2023, while the independent blind validation set (n = 24) consisted of samples grown in a 2023 trial at a single site. The validation set consists of six commercial varieties by four replicates grown at Gerang Gerung, Victoria in sandhill soil over 2023–24 season. While the environment and geographical information of the validation set is not represented in the calibration set, the six commercially released varieties are represented.
For the calibration set, spectra of the wholegrain and grist, the NIR calibration models were constructed using modified partial least squares (mPLS) regression with the following parameters: wavelength region: 1130–2450 nm; scatter: SNV and detrend; derivative: 1, gap: 4; smooth: 4; smooth 2: 1. The modified partial least squares calibration for moisture and protein used six spectral groups, with each group optimized individually for the number of PLS terms. For protein, the average number of mPLS terms selected during cross-validation across all groups was four, while for moisture it was six. A total of six PLS components were used in the final calibration models. The performance of each model was tested using the NIR prediction for the validation set. For calibration construction NIR spectra was collected in triplicate, averaged into a single spectra and assigned a constituent reference value.
Statistical analysis
Two-way ANOVA analysis was conducted using Genstat 24th edition (v24.1.0.1041, VSNi, Hemel Hempstead, UK) to compare the performance of both applications (wholegrain and grist) for each constituent (protein and moisture), using the regression residuals (reference – predicted) for each sample (n = 24).
Results & discussion
Moisture
PLS regression statistics for standard equation and validation set.
N – sample size.
SEC – standard error of calibration.
R2, r2 – coefficient of determination.
SECV – standard error of cross validation.
1-VR – variance subtracted from one.
SEP – standard error of prediction.
RPD – residual prediction deviation (ratio of prediction to deviation).

Calibration plots of NIR-predicted triplicate means versus reference values for lupin wholegrain and ground grist training sets, (a) protein% (w/w) dry basis and (b) moisture% (w/w) content.

Validation plots of reference values versus NIR-predicted triplicate means for lupin wholegrain and ground grist independent validation sets, (a) protein% (w/w) dry basis and (b) moisture% (w/w) content.
Protein
Two-way ANOVA analysis between NIR prediction method (wholegrain and grist) residuals across six lupin varieties (n = 24).
d.f – degrees of freedom.
s.s. – sum of squares.
m.s. – mean square.
v.r. – variance ratio.
F pr. – F probability/p-value.
Insights & implications
Overall, both models performed well in predicting moisture and protein content with acceptable residual errors. There was little difference between grist (protein, R2 = 0.94) and wholegrain (protein, R2 = 0.92) calibrations in lupin (Table 1). This finding is consistent with previous NIR research findings that demonstrated for major constituents there were little difference in R2 and error (SEP) between ground and wholegrain malt calibrations. 17 Due to the small difference between these models, it can be inferred that the lupin hull did not adversely impact the prediction of protein or moisture in the wholegrain lupins.
In a practical industry-focussed setting external to laboratory research purposes, such as on-farm sensing or grain receival, it is difficult to justify additional labour to grind and process lupins for a minimal improvement in model predictions. Another benefit of wholegrain scanning is that even low volumes of seed can be predicted non-destructively, which is useful for screening breeding lines. The differences between grist and wholegrain predictions could be viewed as effectively negligible when compared with the resources spent to prepare large quantities of grist samples for scanning.
To determine the impact of the lupin hull on NIR predictions, this study focused on evaluating the prediction accuracy and precision for major constituents, protein and moisture, which comprise the majority of the composition. Further research could primarily investigate how well the grist and wholegrain NIR models perform in predicting other, minor constituents that make up a smaller portion of the cotyledon composition. To extend upon these research findings, investigation could be conducted into how fibre, lipoxygenase, protein composition, alkaloids and phenolic compounds predict through the lupin hull for wholegrain NIR. For future work in this area, it would be ideal to expand the number of varieties of sweet lupin and include a range of growing environments from Western Australia, South Australia, Victoria and New South Wales, as different varieties will not only have different compositions, but may also have different hull thickness levels depending on the growing conditions and soil types.
Conclusion
NIR models were constructed and found to be feasible for predicting protein and moisture content in both lupin wholegrain and grist applications. Both constituents exhibited strong correlations in both models and are indicative that the intact lupin hull did not negatively impact predictions for moisture or protein content. While the grist model displayed a better predictive performance, the accessibility of wholegrain scanning would likely be more suitable in agricultural applications such as screening breeding lines.
Footnotes
Acknowledgements
This research was funded as part of the “DEE2312-002RTX Capitalising on pulse protein food and feed opportunities domestically and internationally” project, a collaboration between Agriculture Victoria and the Australian Grains Research and Development Corporation (GRDC). The authors wish to acknowledge the Crop Quality laboratory at the Horsham SmartFarm, Dr Jason Brand and Dr Audrey Delahunty for their support in analysing and growing the trials.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by Agriculture Victoria and the Australian Grains Research and Development Corporation (GRDC) through project DEE2312-002RTX Capitalising on pulse protein food and feed opportunities domestically and internationally.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Appendix
NIR average predictive values for moisture and protein across 24 blind validation lupin samples. Reference testing was conducted in duplicate and reported as average. NIR values are displayed as mean of triplicate scans ± standard deviation (SD) reported to two decimal places to better reflect measurement variation. WG refers to wholegrain. Gr refers to grist (ground lupin). aThe Lab Protein & Lab Moisture reference values have a standard error (SET) of ±0.3% and ±0.2% respectively.
Sample ID
Variety
Lab protein% (w/w) drybasis
a
NIR_WG protein% (w/w) drybasis
NIR_Gr protein% (w/w) drybasis
Lab moisture% (w/w)
a
NIR_WG moisture% (w/w)
NIR_Gr moisture% (w/w)
1
Uniharvest
34.3
34.5 ± 0.27
33.9 ± 0.34
11.7
11.5 ± 0.05
11.8 ± 0.03
2
Coyote
31.1
31.1 ± 0.38
31.0 ± 0.04
9.9
9.9 ± 0.02
9.9 ± 0.04
3
PBA Bateman
28.6
29.9 ± 0.35
29.2 ± 0.23
10.1
10.1 ± 0.05
10.1 ± 0.01
4
PBA Gunyidi
31.4
32.0 ± 0.75
31.9 ± 0.40
10.1
10.1 ± 0.02
10.1 ± 0.05
5
Lawler
29.3
29.8 ± 0.20
29.6 ± 0.37
10.2
10.2 ± 0.05
10.2 ± 0.2
6
Mandelup
29.2
29.7 ± 0.41
29.2 ± 0.11
10.8
10.6 ± 0.02
10.8 (SD < 0.01)
7
PBA Bateman
28.8
29.2 ± 0.66
29.1 ± 0.23
10.7
10.6 ± 0.03
10.6 ± 0.02
8
Uniharvest
34.8
36.2 ± 0.38
34.8 ± 0.27
12.0
11.6 ± 0.06
12.1 ± 0.03
9
Coyote
27.9
28.3 ± 0.12
28.4 ± 0.02
10.2
10.2 ± 0.02
10.2 ± 0.02
10
PBA Gunyidi
32.3
32.5 ± 0.68
32.3 ± 0.27
11.9
11.5 ± 0.01
11.9 ± 0.01
11
Mandelup
28.5
29.1 ± 0.26
28.3 ± 0.09
10.7
10.6 ± 0.02
10.7 ± 0.02
12
Lawler
31.1
31.7 ± 0.31
30.7 ± 0.43
10.5
10.4 ± 0.02
10.5 ± 0.04
13
Mandelup
30.2
31.2 ± 0.61
29.9 ± 0.13
10.7
10.6 ± 0.06
10.8 ± 0.03
14
Uniharvest
35.2
37.7 ± 0.80
34.6 ± 0.33
10.2
10.2 ± 0.02
10.2 ± 0.01
15
Lawler
31.4
30.4 ± 0.20
30.7 ± 0.20
10.3
10.2 ± 0.02
10.2 ± 0.03
16
Coyote
29.9
29.8 ± 0.38
29.5 ± 0.17
10.3
10.4 ± 0.02
10.3 ± 0.02
17
PBA Gunyidi
30.9
31.3 ± 0.45
31.0 ± 0.10
9.9
9.9 ± 0.01
9.9 ± 0.04
18
PBA Bateman
30.0
30.0 ± 0.63
29.3 ± 0.25
12.0
11.7 ± 0.03
12.0 ± 0.11
19
Coyote
30.8
30.0 ± 0.53
29.9 ± 0.20
11.0
10.9 ± 0.02
10.9 ± 0.05
20
Lawler
29.9
31.6 ± 0.62
30.5 ± 0.61
10.4
10.3 ± 0.02
10.4 ± 0.08
21
Uniharvest
36.9
37.1 ± 0.72
36.7 ± 0.14
9.7
9.8 ± 0.04
9.7 ± 0.04
22
PBA Bateman
30.3
31.2 ± 0.44
30.7 ± 0.13
10.5
10.4 ± 0.01
10.5 ± 0.07
23
PBA Gunyidi
31.1
32.4 ± 0.73
31.5 ± 0.42
11.5
11.3 ± 0.05
11.5 ± 0.04
24
Mandelup
29.6
29.8 ± 0.54
29.3 ± 0.17
11.9
11.5 ± 0.06
11.9 ± 0.03
Mean
31.0
31.5
30.9
10.7
10.6
10.7
Range
27.9–36.9
28.3–37.7
28.3–36.7
9.7–12
9.8–11.7
9.7–12.1
Pooled repeatability (SD)
0.51
0.27
0.03
0.05
