Hyperspectral leaf reflectance as proxy for photosynthetic capacities: An ensemble approach based on multiple machine learning algorithms
Global agriculture production is challenged by increasing demands from rising population and a changing climate, which may be alleviated through development of genetically improved crop cultivars. Research into increasing photosynthetic energy conversion efficiency has proposed many strategies to improve production but have yet to yield real-world solutions, largely because of a phenotyping bottleneck. Partial least squares regression (PLSR) is a statistical technique that is increasingly used to relate hyperspectral reflectance to key photosynthetic parameters associated with carbon uptake (Vc,max) and conversion of light energy (Jmax) to alleviate this bottleneck. However, its performance varies significantly across different plant species, regions, and growth environments. Thus, to cope with the heterogenous performances of PLSR, this study aims to develop a new approach to estimate photosynthetic parameters. A framework was developed that combines six machine learning algorithms, including artificial neural network (ANN), support vector machine (SVM), least absolute shrinkage and selection operator (LASSO), random forest (RF), Gaussian process (GP), and PLSR to optimize high-throughput analysis of the two photosynthetic parameters. Six tobacco genotypes, including both transgenic and wild-type lines, with a range of photosynthetic capacities were used to test the framework. Leaf reflectance were measured from 400-2500 nm using a high-spectral-resolution spectroradiometer. Corresponding photosynthesis vs. CO2 concentration response curves were measured for each leaf using a leaf gas-exchange system. Results suggested that the mean R2 42 value of the six regression techniques for predicting Vc,max (Jmax) ranged from 0.60 (0.45) to 0.65 (0.56) with the mean RMSE value varying from 47.1 (40.1) to 54.0 (44.7) μmol m-2 s-1 Regression stacking for Vc,max (Jmax) performed better than the individual regression techniques with increases in R2 45 of 0.1 (0.08) and decreases in RMSE by 4.1 (6.6) μmol m-2 s-1 46 , equal to 8% (15%) reduction in RMSE. Better predictive performance of the regression stacking is likely attributed to the varying coefficients (or weights) in the level-2 model (the LASSO model) and the diverse ability of each individual regression technique to utilize spectral information for the best modeling performance. Further refinements can be made apply this stacked regression technique to other plant phenotypic traits.