Light-Invariant Spectral Autoencoder (LISA)
- The paper introduces LISA, a domain-adversarial autoencoder that learns illumination-invariant features from raw hyperspectral grape data for quality prediction.
- It jointly optimizes reconstruction, regression, and domain discrimination losses, yielding over 20% improvement in cross-domain robustness.
- Integrated into an IoT-enabled vineyard pipeline, LISA enables real-time, non-destructive, spatially resolved mapping of grape yield and quality.
to=arxiv_search.search ส่งเงินบาทไทย 大发快三豹子on бжьара 天天彩票是 code 神彩争霸_json {"2query2 OR \2"In-Field Mapping of Grape Yield and Quality with Illumination-Invariant Deep Learning\"", "max_results": 5, "sort_by": "relevance"}】【:】【“】【assistant to=arxiv_search.search RTLU code 全民彩票天天送json {"2query2 hyperspectral illumination invariant autoencoder grapes", "max_results": 2(Cornelissen et al., 6 Oct 2025) OR \2query2, "sort_by": "relevance"}Japgollyassistant to=arxiv_search.search{Jsii code 凤凰大参考 彩票总代理 to=arxiv_search.search _一本道 彩神争霸是_json {"2query2 hyperspectral grape quality illumination-invariant", "max_results": 2(Cornelissen et al., 6 Oct 2025) OR \2query2, "sort_by": "relevance"} to=arxiv_search.search 开号链接 code 彩神争霸输钱json {"2query2 "max_results": 5, "sort_by": "submittedDate"} to=arxiv_search.search 彩神争霸如何 code 天天中彩票未json {"2query2 Mapping of Grape Yield and Quality with Illumination-Invariant Deep Learning\" (&&&2query2&&&)", "max_results": 2(Cornelissen et al., 6 Oct 2025) OR \2query2, "sort_by": "relevance"} to=arxiv_search.search 彩神争霸邀请码 code 大发快三是不是_json {"2query2 yield quality hyperspectral domain adversarial illumination invariant", "max_results": 2(Cornelissen et al., 6 Oct 2025) OR \2query2, "sort_by": "relevance"} Light-Invariant Spectral Autoencoder (LISA) is a domain-adversarial deep learning framework for hyperspectral quality assessment under variable illumination, introduced within an end-to-end, IoT-enabled robotic system for non-destructive, real-time, and spatially resolved mapping of grape yield and quality in vineyards. In the reported system, LISA addresses the in-field hyperspectral imaging domain shift induced by changing illumination and learns illumination-invariant features directly from uncalibrated data, with the aim of predicting grape quality variables—specifically Brix and Acidity—from hyperspectral patches (&&&2query2&&&).
2(Cornelissen et al., 6 Oct 2025) OR \2. Position within in-field hyperspectral viticulture
LISA is one of two key modules in a broader analytical pipeline for vineyard phenotyping. The full system integrates a high-performance model for grape bunch detection and weight estimation with a hyperspectral quality-assessment framework. Within that division of labor, LISA is the component responsible for estimating quality from hyperspectral imaging (HSI) data under uncontrolled field illumination. The complete pipeline is reported to achieve a recall of 2query2.82 for bunch detection and an PRESERVED_PLACEHOLDER_2query2^ of 2query2.76 for weight prediction, while the LISA module improves quality prediction generalization by over 22query2% compared to the baselines and contributes to the generation of high-resolution, georeferenced data of both grape yield and quality (&&&2query2&&&).
The central problem addressed by LISA is the practical difficulty of in-field HSI when illumination varies across time and acquisition conditions. The reported formulation is explicitly motivated by the gap between controlled artificial lighting and natural sunlight acquired at different times of day. Rather than performing radiometric calibration with white-reference panels, LISA learns an invariant representation from raw data. This design choice is significant because the paper frames radiometric calibration as impractical in the field and uses representation learning to absorb illumination variability into the latent space.
2. Architectural formulation
LISA takes as input a raw hyperspectral patch PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \2. The preprocessing stage applies a first derivative with a Savitzky–Golay filter, but the network itself operates on uncalibrated data. The model backbone is a 2D convolutional autoencoder that treats the 224-band spectrum as channels. Its encoder consists of several stacked Conv-BN-ReLU blocks with kernel size , stride 2(Cornelissen et al., 6 Oct 2025) OR \2, and padding 2(Cornelissen et al., 6 Oct 2025) OR \2, reducing the channel dimension from 224 to a lower . The paper does not specify the exact value of , noting only that it is lower; the consolidated description states that the paper does not list exact filter counts per convolutional layer (&&&2query2&&&).
The encoder preserves spatial extent throughout: the spatial dimensions remain and no pooling is used. The decoder mirrors the encoder with Conv-BN-ReLU blocks and reconstructs . This reconstruction pathway supplies an explicit autoencoding constraint on the latent representation.
Two prediction heads are attached to the learned representation . The task predictor, or quality head, flattens and passes it to a small MLP with two hidden layers and ReLU activations, producing PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \2query2^ for predicted Brix and Acidity. The domain discriminator consists of a gradient-reversal layer (GRL), followed by an MLP with two hidden layers and ReLU activations, and a softmax output PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \2(Cornelissen et al., 6 Oct 2025) OR \2^ over the three acquisition domains: Lab, Field-AM, and Field-PM.
A notable property of this architecture is that it couples reconstruction, regression, and domain discrimination within a shared latent representation. This suggests that the latent variable PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \22^ is intended to remain predictive of grape chemistry while discarding illumination-specific information.
3. Optimization objectives and adversarial mechanism
LISA is trained end-to-end with the total loss
PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \23
where the best reported weighting hyperparameters are PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \24, PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \25, and PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \26 (&&&2query2&&&).
The reconstruction term is
PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \27
described as the Frobenius-norm squared, or MSE over all pixels and bands. The task loss is
PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \28
with PRESERVED_PLACEHOLDER_2(Cornelissen et al., 6 Oct 2025) OR \29 the ground-truth vector of Brix and Acidity and 2query2^ the network prediction. The manifold regularization term is
2(Cornelissen et al., 6 Oct 2025) OR \2^
where 2 if 3 and 4 fall in the same Brix category, and 5 otherwise. According to the reported interpretation, this term encourages samples with similar Brix to lie close in latent space.
The domain-adversarial component is implemented as a two-player min-max game using GRL, without separate alternating updates. The discriminator 6 minimizes
7
the standard cross-entropy over true domain labels 8. Through the GRL, the encoder 9 maximizes that same cross-entropy, equivalently minimizing
2query2^
after the GRL flips the gradient. During each minibatch, samples from all three domains are passed through the shared encoder, the GRL routes 2(Cornelissen et al., 6 Oct 2025) OR \2^ to the discriminator, and on back-propagation the GRL multiplies 2 by 3. The resulting single back-propagation pass updates the encoder to fool the discriminator while updating the discriminator to classify domains correctly.
The ablation results reported in the paper identify the adversarial component as the main driver of domain invariance. This is consistent with the formulation, since the GRL enforces pressure against domain-separable latent structure while retaining the regression and reconstruction constraints.
4. Data acquisition and preprocessing regime
The reported dataset was purpose-built to evaluate robustness across illumination domains. Data were acquired with a Specim FX2(Cornelissen et al., 6 Oct 2025) OR \2query2^ push-broom HSI camera with 4 spatial pixels, 224 bands, and spectral range 42query2query2–2(Cornelissen et al., 6 Oct 2025) OR \2query2query2query2^ nm. The acquisition protocol sampled 62query2^ grape bunches, each imaged under three domains: Lab with controlled halogen lighting, Field-AM under morning sun, and Field-PM under afternoon sun (&&&2query2&&&).
Quality annotations were derived from 362query2^ manually annotated single-berry spectral patches per domain, for a total of 2(Cornelissen et al., 6 Oct 2025) OR \2,2query2 hyperspectral illumination invariant autoencoder grapes2query2^ patches. Ground-truth Brix was measured with a refractometer and acidity by titration. Preprocessing consisted of passing raw Digital Numbers through a Savitzky–Golay first-derivative filter with window length 2(Cornelissen et al., 6 Oct 2025) OR \25 and polynomial order 2, followed by extraction of 5 patches from spatially homogeneous grape regions.
This data design is methodologically important because the three-domain setup turns illumination variation into an explicit domain-generalization problem. Lab, Field-AM, and Field-PM are not merely nuisance conditions; they define the domain labels used by the adversarial discriminator. A plausible implication is that the learned invariance is tied specifically to the acquisition conditions represented in those three domains, rather than to illumination variation in the abstract.
5. Empirical behavior under domain shift
The reported regression metric is the coefficient of determination,
6
For classification of grape versus non-grape, the metric is Overall Accuracy (OA), and for object detection in the yield module the paper uses Recall and mAP@2query2.5:2query2 (&&&2query2&&&).
The core evaluation for LISA is leave-one-domain-out generalization. When trained on Lab+PM and tested on Field-AM, LISA achieves Brix 7, compared with Predictor-Only at 2query2.42(Cornelissen et al., 6 Oct 2025) OR \2, SiLLR-GAN at 2query2.43, PLS at 2query2.25, and LOGSEP at 2query2.2(Cornelissen et al., 6 Oct 2025) OR \27. In the same condition, Acidity 8 is reported as best, and grape OA is 2(Cornelissen et al., 6 Oct 2025) OR \2.2query2query2. When trained on Lab+AM and tested on Field-PM, LISA achieves Brix 9, Acidity 2query2, both reported as best, and grape OA of 2query2.99.
The ablation study isolates the contribution of each loss component. The full model with adversarial learning yields 2(Cornelissen et al., 6 Oct 2025) OR \2, whereas removing 2 reduces performance to 2query2.52query2 Removing 3 gives 4, and removing 5 gives 6. These results support the paper’s conclusion that the adversarial term is the principal mechanism for cross-domain robustness, while manifold regularization and reconstruction provide additional but smaller gains.
The latent-space visualization further supports that interpretation. The paper reports that t-SNE on raw patches clusters by domain, whereas the learned latent representation 7 clusters by Brix and is well mixed across domains. This suggests that the representation is reorganized around task-relevant chemistry rather than acquisition condition.
6. Training regime, deployment characteristics, and interpretive boundaries
The reported optimizer is Adam. The encoder/decoder and task predictor use learning rate 8, while the domain discriminator uses 9. Training runs for 22query2query2^ epochs. The batch size is not specified in the paper. For deployment, the hardware is an Edge PC, specifically an Intel NUC2(Cornelissen et al., 6 Oct 2025) OR \2query2i5FNH, and no GPU is needed for real-time inference; the entire pipeline is reported to run faster than the push-broom acquisition rate on CPU (&&&2query2&&&).
The deployment result matters because LISA is not presented as a purely offline spectral model. It is part of a real-time robotic sensing system operating in the field. In that setting, the ability to work from uncalibrated data and to infer on CPU aligns with the paper’s emphasis on operational practicality.
Several interpretive boundaries are explicit in the reported conclusions. First, LISA addresses radiometric calibration impracticalities by learning invariance directly from uncalibrated data and removes the need for white-reference panels in the field. Second, the remaining gap is not eliminated: cross-domain 2query2^ of 2query2.62 is described as an approximately 22query2% improvement over baselines but still below lab-only performance. This suggests that domain invariance is partial rather than absolute.
A common misconception would be to read LISA as an explicit reflectance-recovery method. The reported formulation instead learns a domain-invariant latent representation for prediction; future work in the paper proposes hybrid physics-data frameworks for more interpretable reflectance recovery. Additional suggested directions are scaling the dataset across varieties, growth stages, and locations; using semi-supervised or advanced data-augmentation methods to reduce annotation effort; and integrating temporal models to exploit sunlight drift. These directions indicate that the authors regard the current three-domain formulation as a strong operational baseline rather than a complete account of in-field hyperspectral variability.