Resolve the reliability gap in few-shot hyperparameter search on Gemma-4
Establish a reliable hyperparameter-selection procedure for soft prompting on Gemma-4-12B-it that allows performance on an internal holdout drawn from the ten-shot support set to predict performance on unseen test images.
References
We report this as an open reliability gap in the search procedure itself, not as evidence against the underlying method: the same $10$-shot budget that lets \S\ref{sec:ablations}'s Qwen3-VL search reliably rank configurations does not, on this second backbone, carry enough signal to always rank them correctly.
— Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models
(2609.11310 - Gare et al., 10 Sep 2026) in Appendix, Section “Gemma-4 Generalisation: Search Methodology and Full Results,” final paragraph