Scaling ethical-preference behavior to larger language models

Determine whether the ethical-preference learning and task-vector behavior observed in Llama-3.2 1B and 3B checkpoints scales to larger language models.

Background

The study evaluates ethical stance adherence and preference reversal only on Llama-3.2 models with 1B and 3B parameters. Although these compact models achieve high accuracy after LoRA or DPO fine-tuning and support task-vector-based preference reversal, the experiments do not establish whether the same behavior persists as model size increases. The authors therefore identify the scaling of these findings to larger models as unresolved.

References

Fourth, our experiments are limited to Llama-3.2 1B and 3B checkpoints, and it remains open whether the same behavior scales to larger models.

— Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models  (2609.21094 - Agarwal et al., 17 Sep 2026) in Section 6, Limitations