LeCropFollow: Vision-Based Crop Navigation
- LeCropFollow is a vision-based navigation framework that replaces explicit crop-row geometry with learned latent representations to handle ambiguities in under-canopy settings.
- It integrates a self-supervised semantic heatmap extractor with TD-MPC2 model-based planning, achieving 93.3% success in traversing an 8.7 m plantation gap.
- The system preserves high-dimensional uncertainty in heatmaps to reduce semantic failures by 2.4× compared to conventional keypoint-based approaches.
Searching arXiv for LeCropFollow and closely related crop-following navigation papers. LeCropFollow is a vision-based under-canopy navigation framework for agricultural robots that replaces explicit crop-row geometry estimation with planning in a learned latent space derived from semantic heatmaps. It targets the failure modes that remain prominent in under-canopy autonomy—irregular planting, missing rows, occlusions, and plantation gaps—by preserving high-dimensional semantic context instead of compressing observations into deterministic spatial references such as keypoints or centerlines. The framework couples a self-supervised semantic heatmap extractor with TD-MPC2, a model-based reinforcement learning planner, and is reported to achieve zero-shot transfer from simplified simulation to physical late-stage corn fields without fine-tuning; in an 8.7 m plantation gap it attains 93.3% success and a 2.4× reduction in semantic failures relative to keypoint-based baselines (Tommaselli et al., 30 Jun 2026).
1. Problem setting and motivation
LeCropFollow addresses under-canopy navigation in GPS-denied, narrow crop corridors, where robots must operate using onboard sensing while coping with severe ambiguity in row structure. The paper identifies unstructured navigational features—especially irregular planting and discontinuities—as the primary failure mode for under-canopy agricultural robots, and reports that about half of autonomy failures in prior work arise from such unstructured features (Tommaselli et al., 30 Jun 2026).
The central critique is directed at conventional geometric pipelines. In typical under-canopy systems, perception estimates an explicit reference such as a row centerline, vanishing point, or a small set of semantic keypoints, and control is then conditioned on that reduced representation. LeCropFollow argues that this reduction is especially brittle in plantation gaps, because the representation remains low-dimensional even when the scene is semantically ambiguous. In that regime, a controller is forced to act on a point estimate that may no longer correspond to meaningful row geometry.
This motivates a representational shift rather than a purely algorithmic refinement of the same geometric abstraction. The framework is designed to operate over the uncompressed semantic heatmap signal, preserving the dispersion of the heatmaps and therefore the uncertainty that would otherwise be discarded by an argmax-based keypoint extraction stage. A plausible implication is that the method is less dependent on the existence of a well-posed geometric corridor at every instant, which is precisely the condition that fails in plantation gaps.
2. Departure from geometric crop-following representations
Earlier under-canopy crop-following systems commonly encode the traversable corridor using three semantic keypoints: a vanishing point, a left intersection or intercept point, and a right intersection or intercept point. In MetaCropFollow and AdaCropFollow, these three points define a triangle corresponding to the free navigable space between crop rows; both works retain the keypoint abstraction and then improve generalization through meta-learning or self-supervised online adaptation, respectively (Woehrle et al., 2024, Sivakumar et al., 2024).
LeCropFollow departs from that design at the representation level. Its perception stage produces a semantic heatmap tensor
whose three channels correspond to the Vanishing Point, Left Row, and Right Row. Instead of taking the argmax of each channel to obtain keypoints, the method keeps the full tensor and feeds it to a latent encoder together with the previous action:
The inclusion of the previous action is justified by the observation that visual data alone cannot infer the current velocity and steering state (Tommaselli et al., 30 Jun 2026).
This design is conceptually distinct from explicit geometry estimation. Heatmaps can be diffuse in occluded or discontinuous regions, and that dispersion itself is treated as useful information. The paper’s claim is not merely that a learned representation is richer than a hand-crafted one, but that uncertainty and semantic context are already encoded in the heatmap and are lost when the system is forced into a deterministic geometric bottleneck. In LeCropFollow, planning is conditioned on the latent representation of that heatmap manifold rather than on explicit row geometry.
3. System architecture and planning formulation
The framework has three principal stages: perception, latent world-model learning, and online planning. Perception uses a frozen self-supervised backbone, specifically RowFollowNet from CropFollow++, to process a monocular RGB image
into the semantic heatmap tensor . The downstream action space follows unicycle kinematics, with
where is linear velocity and is angular velocity. Policy outputs are normalized to and affine-mapped to m/s and rad/s. Before actuation, the command is smoothed by a first-order exponential filter:
0
This is intended to attenuate jitter (Tommaselli et al., 30 Jun 2026).
The model-based reinforcement learning core is TD-MPC2. The learned components are the latent encoder 1, prior policy 2, dynamics model 3, reward model 4, and value function 5. The task is written as local trajectory optimization in a POMDP
6
with the control decision defined by a finite-horizon objective:
7
Future latent states are rolled out according to
8
and only the first action of the optimized sequence is executed before replanning in receding-horizon fashion (Tommaselli et al., 30 Jun 2026).
Online planning is performed with Model Predictive Path Integral control over the latent manifold. At each step, the system encodes the current observation, samples perturbed action sequences from a Gaussian centered on the policy prior, rolls each candidate through latent dynamics over horizon 9, scores them using the planned rewards plus terminal value, and executes the reward-weighted average of the top-0 elite sequences. The paper’s claim is that this latent rollout mechanism allows the robot to “look past” local discontinuities such as plantation gaps rather than reacting myopically to a noisy instantaneous geometric estimate.
The reward is described as a logistic gating reward from CaRL. Its components are
1
2
and
3
with 4, 5, 6, and 7. The intended effect is that fast motion is rewarded only when the platform remains stable and collision-free (Tommaselli et al., 30 Jun 2026).
4. Training regime, simulation design, and deployment
Training is conducted in Gazebo using simplified geometric primitives rather than photorealistic plant models. Crop obstacles are represented as cylinders, row spacing is 0.75 m, and colors are randomized with 80% green, 10% red, and 10% blue. Episodes terminate on collision, on reaching a 120 m forward-distance threshold, or after a 500-step budget, with a grace window of roughly ten steps at the start to avoid spurious resets during spawn transients (Tommaselli et al., 30 Jun 2026).
The simulation strategy is deliberately minimal. The paper argues that the simplified visuals constitute a lower-bound out-of-distribution case for the frozen heatmap backbone: although rendered observations are crude, the downstream heatmap signal is sufficient for learning. This suggests that the sim-to-real transfer claim rests less on visual realism than on the invariance of the semantic heatmap representation across domains.
Implementation is compact but fully online. The learned system is a 5M-parameter MLP with Mish activations comprising 8, 9, 0, 1, and 2. Training runs for 120k steps, taking 11.4 hours on an NVIDIA RTX A2000, and deployment is performed on an NVIDIA Jetson Orin Nano. Planner and policy run asynchronously at 20 Hz. Reported hyperparameters include MLP width 512, depth 3, batch size 512, SimNorm dimension 8, five Q-functions, learning rate 3, buffer size 4, seed steps 5500, MPPI iterations 3, 5 samples, temperature 0.50, 6 elites, 16 policy trajectories, MPPI 7, MPPI 8, and policy prior coefficient 0.20 (Tommaselli et al., 30 Jun 2026).
Field experiments use a TerraSentia skid-steer robot equipped with wheel encoders, a 6-DoF IMU, and a ZED 2i camera using the right monocular stream. A Livox Mid-360 LiDAR is present only for the CROW baseline. An important methodological point is that LeCropFollow does not require odometry, mapping, or LiDAR at inference, whereas both CropFollow++ and CROW use DLIO from LiDAR and IMU for ego-motion estimation in the comparative experiments (Tommaselli et al., 30 Jun 2026).
5. Empirical performance and failure analysis
The empirical evaluation is conducted in two late-season corn settings, flowering stage and harvested stage, which share row morphology but differ photometrically and in occlusion conditions. In 12 flowering-stage runs, LeCropFollow records an average of 5.1 collisions and an average maximum distance without collision of 27.9 m; the comparative flowering-stage table reports 5.3 ± 1.6 collisions and 38.9 m for CropFollow++, 6.5 ± 1.4 and 37.9 m for CROW, and 5.1 ± 1.8 and 38.9 m for LeCropFollow. In the harvested stage, LeCropFollow attains 4.8 ± 1.5 collisions and 28.4 ± 5.7 m maximum collision-free distance. Mann-Whitney 9 tests are reported as showing no significant difference between flowering and harvested distributions for collisions or distance, which the paper interprets as evidence that the policy is not overfit to a single field appearance (Tommaselli et al., 30 Jun 2026).
The principal result concerns plantation gaps. In a row containing an 8.7 m gap, CropFollow++ achieves 6.7% success, CROW 53.3%, and LeCropFollow 93.3%, corresponding to 14 successful traversals out of 15. The paper attributes this advantage to the planner’s behavior under high uncertainty: rather than blindly tracking noisy keypoint predictions, LeCropFollow makes small angular corrections and maintains a smooth heading until recoverable row structure reappears (Tommaselli et al., 30 Jun 2026).
Failure analysis distinguishes semantic failures from physical failures. Semantic failures comprise Perception Error and Occlusion; physical failures comprise Actuation, Bad Start, and Terrain. LeCropFollow is reported to reduce semantic failures by 2.4×, specifically 29 versus 70 relative to baselines. This is the paper’s strongest evidence that latent-space planning improves robustness to ambiguity in the observation model rather than merely improving low-level control (Tommaselli et al., 30 Jun 2026).
The ablation study reinforces that interpretation. The full system is denoted L3; L2 reduces the planning horizon to essentially one-step planning; L1 uses only the policy prior without the planner; G1, G2, and G3 are geometric variants. The reported findings are that L3 performs best, removing the planner degrades performance substantially, shortening the horizon collapses performance close to CropFollow++, and Gaussian sampling alone is insufficient. The improvement is therefore attributed to the combination of latent heatmap representation and model-based planning, not to perception alone.
6. Relation to adjacent work, limitations, and significance
Within under-canopy navigation, LeCropFollow occupies a distinct position relative to both keypoint-based learning systems and hybrid farm-scale autonomy frameworks. MetaCropFollow retains the three-keypoint corridor representation and uses MAML, MAML++, and ANIL to enable few-shot adaptation across seasonal domains, adapting to a target day from 0 labeled images and showing strong gains under seasonal shift (Woehrle et al., 2024). AdaCropFollow likewise preserves the semantic keypoint representation but introduces a self-supervised online adaptation loop driven by stereo zero-disparity for vanishing points and geometric-prior-guided pseudo-labels for intercept points, updating 111,360 parameters onboard the robot computer (Sivakumar et al., 2024). Both methods treat domain shift as the central obstacle while maintaining explicit geometric abstraction.
LeCropFollow instead treats the abstraction itself as a source of brittleness in unstructured terrain. In that sense, it is not an adaptation layer for a keypoint-based pipeline but a replacement for the keypoint-to-geometry stage. This does not make the earlier methods obsolete; rather, it identifies a different failure mode. A plausible implication is that keypoint adaptation and latent planning address complementary axes of robustness: the former improves cross-domain perception under recognizable row structure, while the latter improves action selection when row structure becomes semantically ambiguous.
At the system level, CropNav extends row-following into a full-field navigation framework that switches between LiDAR row following and GNSS waypoint navigation, detects failures through a traction-related coefficient 1, and recovers autonomously, reporting about 750 m per intervention over GNSS-based navigation and 500 m over row-following navigation (Gasparino et al., 2024). LeCropFollow is narrower in scope: it is an under-canopy perception-and-planning framework rather than a whole-farm supervisory architecture. Its empirical strength lies in ambiguous in-row traversal, especially plantation gaps, not in headland transitions or multi-modal navigation.
The paper is explicit about limitations. Physical failures remain, including terrain irregularities, actuator limits, and heading spikes. The evaluation is focused on gap geometry and row-following-style scenarios, and the authors note that better low-level closed-loop control could improve robustness. They also suggest that additional sensing modalities such as depth could help (Tommaselli et al., 30 Jun 2026).
The significance of LeCropFollow is therefore representational as much as empirical. It replaces the conventional pipeline detect keypoint 2 estimate geometry 3 control
with RGB image 4 semantic heatmap 5 latent encoding 6 world-model planning 7 action. Its reported zero-shot transfer from simplified simulation, parity with strong baselines in standard late-stage rows, and substantial gains in plantation gaps collectively suggest that latent planning is a viable alternative to explicit geometric estimation in heterogeneous under-canopy agricultural environments (Tommaselli et al., 30 Jun 2026).