- The paper shows that single-GPU VLA fine-tuning can produce a silent 29-percentage-point seed lottery on LIBERO-Object, with 12 of 13 baseline seeds succeeding but one collapsing to 65.2% success.
- The paper explains that weight-level methods such as EWC and L2 fail to constrain output-active directions under frozen encoders, while patch-level VICReg, Dropout, and reduced learning rates eliminate catastrophic seeds across 21 combined runs.
- The paper finds that output-level methods preserve strong average performance while reducing tail risk, with LowLR reaching 94.5% Object success, VICReg reaching 93.6% Goal success, and Dropout achieving 95.4% Spatial success.
The seed lottery phenomenon
This paper documents a reproducibility failure in single-GPU fine-tuning of vision-language-action (VLA) models and proposes a mechanistic explanation together with three practical remedies. The setting is VLA-JEPA, initialized from a checkpoint pretrained on LIBERO at 8×A100 scale, and fine-tuned on a single RTX 5090 with both encoders frozen. Running identical fine-tuning code with thirteen different random seeds on LIBERO-Object produces a starkly bimodal distribution: twelve seeds reach 91–94% success rate, while one collapses to 65.2%. The authors term this the seed lottery — a 29 pp spread with no training-time warning signal, no loss anomaly, and no way to predict which seed will fail.
The lottery is suite-dependent. LIBERO-Spatial shows essentially none (Baseline std 1.0%, n=9), while LIBERO-Object is the primary affected suite. The authors attribute this asymmetry to task structure: Spatial tasks require positional precision, where a collapsed action predictor that defaults to a mean grasp still succeeds on geometrically average objects; Object tasks require visual discrimination between similar items (e.g., cream cheese versus butter at 6.7 cm separation), which collapse destroys. This hypothesis is supported but not fully established — the paper explicitly notes that the precise mechanism of suite-dependent lottery intensity remains open.
Mechanism: Jacobian null-space under frozen encoders
The central theoretical claim is that output collapse — the trainable ACPredictor producing nearly observation-independent outputs — occurs in directions of parameter space that weight-level regularizers cannot see. Formally, performance depends on ∥JΔθ∥ rather than ∥Δθ∥, where J=∂gθ/∂θ is the Jacobian of the trainable head. With a frozen encoder, the input distribution to gθ is stationary, making J rank-deficient with dim(null(J))≫rank(J). L2 and EWC penalize ∥Δθ∥ uniformly, spending their budget on output-inert null-space directions while leaving the output-active subspace unconstrained.
The strongest empirical support is what the authors call the EWC paradox: EWC achieves the smallest weight drift of any fine-tuned method (∥Δθ∥2=7.8, closest to the pretrained checkpoint) yet collapses to 65.8% Object SR — worse than the Baseline mean of 91.0%. Conversely, VICReg permits five times more drift (40.2) yet achieves 92.7%, and L2 with eighteen times more drift (140.4) also collapses at 62.0%. Weight proximity to the pretrained checkpoint therefore does not predict task performance. Notably, this observation leaves dim(null(J)) unmeasured, so the null-space account is consistent with the data rather than directly verified. Three falsifiable predictions follow from the theory — encoder unfreezing narrows the gap, LowLR succeeds via early-phase trajectory avoidance rather than weight anchoring, and VICReg+L2 does not beat VICReg alone — of which predictions (i) and (ii) are reported as empirically confirmed.
Method: patch-level VICReg at batch size one
The architectural substrate follows VLA-JEPA: a frozen V-JEPA2 ViT-L/16 encodes video frames into patch embeddings, a frozen Qwen3-VL-2B encodes image and language instruction into action tokens, and two trainable modules — a 12-layer ACPredictor trained with an L1 world-model loss against stop-gradient targets, and a 16-layer DiT flow-matching action head producing 7-DOF chunks at horizon ∥JΔθ∥0 — are fine-tuned jointly.
The methodological obstacle for applying VICReg (Bardes et al., 2021) here is that standard VICReg requires large batches, whereas single-GPU fine-tuning operates at micro-batch ∥JΔθ∥1. The paper's solution reshapes predicted latents along the spatial-temporal patch dimension, yielding ∥JΔθ∥2 rows treated as the sample axis for variance and covariance losses (∥JΔθ∥3, ∥JΔθ∥4 following DynaMo). The invariance term is deliberately omitted, since predictions should vary with observations. Stability at ∥JΔθ∥5 is validated empirically through tight cross-seed statistics and near-deterministic drift (∥JΔθ∥6 across three seeds), which the authors interpret as the fingerprint of a unique attractor imposed by the output constraint.
Two complementary instantiations test whether the taxonomy, rather than VICReg specifically, carries the finding: Dropout (∥JΔθ∥7) applied to both trainable modules, and LowLR, which halves both learning rates with zero code changes. LowLR is classified as output-level despite its weight-space implementation, on the argument that it slows early-phase traversal into collapse basins before the action loss can correct course; its drift (~10.1), comparable to EWC's, supports the claim that drift magnitude alone does not determine collapse.
Main results
Across 7 methods × up to 13 seeds × 3 LIBERO suites, the output/weight taxonomy separates methods cleanly on LIBERO-Object:
| Method |
Class |
∥JΔθ∥8 |
Spatial |
Object |
Goal |
| Pretrained reference |
— |
1 |
96.4 ± 1.7 |
78.0 |
82.0 |
| Baseline |
None |
9–13 |
95.1 ± 1.0 |
91.0 ± 7.5 |
90.5 ± 0.9 |
| Dropout |
Output |
3–4 |
95.4 ± 0.4 |
91.8 ± 0.8 |
90.8 ± 2.2 |
| VICReg |
Output |
9–12 |
94.9 ± 0.5 |
92.7 ± 1.4 |
93.6 ± 1.4 |
| LowLR |
Output |
5 (Obj.) |
94.6 ± 2.0 |
94.5 ± 0.7 |
92.8 |
| EWC (identity Fisher) |
Weight |
4 (Obj.) |
86.8 ± 3.0 |
65.8 |
67.4 |
| L2 |
Weight |
1 |
87.2 ± 2.9 |
62.0 |
66.0 |
All output-level methods record zero catastrophic seeds across 21 combined runs (∥JΔθ∥9, ∥Δθ∥0; Levene robust ∥Δθ∥1), compressing the 29 pp Baseline span to ≤4 pp. All three surpass the pretrained Object reference of 78.0%.
An important framing point: the contribution is tail-risk elimination, not mean improvement. Excluding the collapsed seed, the Baseline reaches 92.4% ± 1.3% (∥Δθ∥2), statistically indistinguishable from VICReg's 92.7% ± 1.4% (∥Δθ∥3, ∥Δθ∥4). Practitioners who can afford multi-seed evaluation may obtain comparable means without any regularizer; those who cannot afford even one silent deployment failure benefit from output-level constraints.
Collapse diagnostics localize the signal precisely. Latent-level cosine similarity (~0.998) is universal across all conditions and uninformative for predicting success; only action-level diversity discriminates, with VICReg a qualitative outlier in mean pairwise cosine similarity (0.571, a 0.073 gap below the nearest method, robust at ∥Δθ∥5 and ∥Δθ∥6). Consequently, predicting collapse from encoder representations is not possible — the diagnostic signal exists exclusively at the predictor output.
The per-task analysis identifies LIBERO-Object Task 1 (cream cheese pick-and-place) as a diagnostic canary: collapsed weight-level seeds score 0–2% there, while every output-level seed scores 82–100%. Non-collapsed EWC seeds reach 86–90% on the same task, confirming the failure is seed-lottery collapse rather than architectural incapacity. Task 4 (ketchup) is trivial for all methods (98–100%) because its cylindrical geometry removes the discriminative requirement. On LIBERO-Spatial, per-task gains reach +32 pp for VICReg over Baseline, though Task 5 shows an −8 pp regression, suggesting the variance floor slightly penalizes low-variance vertical placement actions.
A ∥Δθ∥7 sweep spanning 20× (0.1–2.0) shows ≤5 pp spread on any suite, indicating VICReg is not fragile to hyperparameter choice. Offline metrics reinforce a cautionary note: both from-scratch runs improve offline L1 while achieving 0% closed-loop SR, confirming that closed-loop evaluation is the only reliable quality measure.
Limitations
The evidence base has several acknowledged constraints. Results are confined to VLA-JEPA on LIBERO simulation; cross-architecture validation (SmolVLA variants) is ongoing and real-robot validation is planned but absent. Seed counts for Dropout (∥Δθ∥8–4) and LowLR (∥Δθ∥9) are insufficient to individually confirm lottery elimination with statistical confidence; EWC uses identity Fisher rather than empirical Fisher; L2 has a single seed. The null-space dimension is never measured, leaving Observation 1 formally unverified. A data-mixture confound — joint four-suite training possibly diluting Object discrimination — cannot be ruled out. Finally, the mechanism explaining why Spatial tasks escape the lottery remains hypothetical.
Conclusion
The paper identifies a concrete, quantified reproducibility hazard in frozen-encoder VLA fine-tuning — roughly a 1-in-13 chance of undetectable catastrophic failure per run on discrimination-heavy suites — and demonstrates that it is eliminated by constraining outputs rather than weights. The practical prescriptions are ordered by cost: halving learning rates (zero code changes, best Object SR at 94.5%), adding Dropout (one line, best Spatial SR), or patch-dimension VICReg (~30 lines, best Goal SR and tightest cross-seed consistency). The sharpest takeaway for practitioners is that weight-proximity regularization is not merely ineffective but counterproductive under frozen encoders, and that the cheapest intervention is a single optimizer configuration change.