Behavior of larger-than-8B vision-language models

Determine whether vision-language models with 32 billion or more parameters behave in the same way as the 2–8B edge-deployable vision-language models evaluated for species identification, including their taxonomic knowledge and performance characteristics.

Background

The study deliberately evaluates vision-LLMs in the 2–8B parameter range because this size class is relevant to on-device camera-trap deployment. The authors explicitly delimit their conclusions to this range and do not evaluate substantially larger models. Consequently, whether the reported behavior extends to 32B or larger models remains unresolved.

References

Whether a 32B or larger model behaves the same way is a genuinely open question we do not address here.

Can Edge-Deployable Vision-Language Models Identify Species?  (2609.11916 - Zhou et al., 10 Sep 2026) in Section 1, Introduction